# Extcord: full catalog

> Everything Extcord can do, in one page: 43 ops with their arguments and a worked example, the recipes, and how to chain ops. The short version is /llms.txt.

## Common jobs in one call

Over MCP these are tools; over HTTP they're recipes (`POST /recipes/{name}/run` with `{"params": {...}}`) or ops (`POST /ops/{id}`).

- **Read a web page, only the part you need** — MCP `read_page(url, view)`; HTTP recipes `page_to_markdown`, `page_outline`, `page_links`, `page_tables`, `find_on_page`.
- **Exact arithmetic** — MCP `calculate_expression(expr, vars)`; HTTP `POST /ops/calc`.
- **Dates, times and time zones** — MCP `calculate_date(action, ...)`; HTTP ops `now`, `date_add`, `date_diff`, `date_info`.
- **Package facts** (PyPI, npm, crates.io) — MCP `read_package_info(ecosystem, name)`; HTTP recipes `pypi_package`, `npm_package`, `crate_info`.

## Chaining ops

A pipeline is a list of steps `{op, input?, args?, as?}`. Refer to earlier results as `$0`, `$1`, `$name` (a step with `"as": "name"`) or a path like `$1[0]` or `$2.items[*].name`; refs work inside args too. Only the final result comes back; the rest stays on the server as handles. Write a literal string starting with `$` as `$$`.

```json
{"steps": [{"op": "fetch", "args": {"url": "https://example.com/pricing"}, "as": "page"},
           {"op": "extract_tables", "input": "$page", "args": {"index": 0}},
           {"op": "to_markdown", "input": "$1"}]}
```

## Operations

### data

#### json_parse

Parse JSON text into data. lenient=true also handles code fences, chatter around the JSON and trailing commas (typical LLM output).

Input: text. Returns: depends on the data.
Args:
- `lenient` (boolean; default false): Repair common problems before parsing.
Example: `{"input": "Sure! Here it is:\n```json\n{\"name\": \"Ada\", \"tags\": [\"x\", \"y\",],}\n```", "args": {"lenient": true}}` → `{"name": "Ada", "tags": ["x", "y"]}`

#### get

Read a value out of JSON by path, e.g. data.items[0].name or items[*].price.

Input: json, list, table. Returns: depends on the data.
Args:
- `path` (string; required): Path into the data: dots for keys, [n] for index (-1 = last), [*] for every item, ["odd key"] for unusual keys.
Example: `{"input": {"data": {"items": [{"name": "A", "price": 3}, {"name": "B", "price": 5}]}}, "args": {"path": "data.items[*].price"}}` → `[3, 5]`

#### filter

Keep the rows (or list items) that match conditions like {field, op, value}. Numbers in strings ('$1,200') compare numerically.

Input: list, table. Returns: same type as input.
Args:
- `where` (array; required): Conditions: [{"field": "price", "op": "lt", "value": 20}]. Ops: eq, ne, gt, gte, lt, lte, contains, not_contains, starts_with, ends_with, in, exists, matches. Omit 'field' to test list items themselves.
- `mode` (string; default "all"; one of all, any): all: every condition must match. any: at least one.
- `case_sensitive` (boolean; default false): Case-sensitive string comparisons.
Example: `{"input": [{"plan": "Free", "price": "$0"}, {"plan": "Pro", "price": "$20"}, {"plan": "Team", "price": "$45"}], "args": {"where": [{"field": "price", "op": "gt", "value": 10}, {"field": "price", "op": "lt", "value": 40}…` → `[{"plan": "Pro", "price": "$20"}]`

#### pick

Pick several fields out of JSON into one small object, by path, under names you choose. The cheapest way to read a few facts from a large API response.

Input: json, list, table. Returns: json.
Args:
- `fields` (object; required): Output name -> path, e.g. {"version": "info.version", "python": "info.requires_python"}.
- `missing` (string; default "null"; one of null, error): When a path doesn't exist: null (put null, the default) or error (stop with a message saying which).
Example: `{"input": {"info": {"version": "2.32.3", "requires_python": ">=3.8", "license": "Apache-2.0"}, "urls": [{"upload_time": "2024-05-29"}]}, "args": {"fields": {"version": "info.version", "python": "info.requires_python", "…` → `{"version": "2.32.3", "python": ">=3.8", "released": "2024-05-29", "homepage": null}`

#### pluck

Take one column from rows as a list, or several columns as a smaller table.

Input: table, list. Returns: depends on the data.
Args:
- `field` (string): One field (or path) to take; returns a list.
- `fields` (array): Several fields to keep; returns a table with only these columns.
Example: `{"input": [{"a": 1, "b": 2}, {"a": 3, "b": 4}], "args": {"field": "a"}}` → `[1, 3]`

#### sort

Sort rows or list items, numerically when values look like numbers.

Input: list, table. Returns: same type as input.
Args:
- `by` (string): Field to sort rows by. Omit for a list of plain values.
- `desc` (boolean; default false): Largest first.
Example: `{"input": [{"p": "$20"}, {"p": "$5"}, {"p": "$100"}], "args": {"by": "p", "desc": true}}` → `[{"p": "$100"}, {"p": "$20"}, {"p": "$5"}]`

#### unique

Remove duplicates from a list or rows, keeping first occurrences.

Input: list, table. Returns: same type as input.
Args:
- `by` (string): For rows: field that defines a duplicate. Omit to compare whole rows.
Example: `{"input": ["a", "b", "a"]}` → `["a", "b"]`

#### count_by

Count how often each value of a field appears; returns rows of {value, count}, most common first.

Input: list, table. Returns: table.
Args:
- `field` (string): Field to group by. Omit for a list of plain values.
Example: `{"input": [{"s": "open"}, {"s": "closed"}, {"s": "open"}], "args": {"field": "s"}}` → `[{"value": "open", "count": 2}, {"value": "closed", "count": 1}]`

#### stats

Summary numbers for a numeric field or list: count, sum, min, max, mean, median.

Input: list, table. Returns: json.
Args:
- `field` (string): Numeric field in rows. Omit for a list of numbers.
Example: `{"input": ["$10", "$20", "$60"]}` → `{"count": 3, "sum": 90, "min": 10, "max": 60, "mean": 30, "median": 20, "skipped": 0}`
Note: Strings like '$1,200' or '15%' are read as numbers; anything non-numeric is skipped and counted in 'skipped'.

### format

#### to_markdown

Render rows as a markdown table, a list as bullets, or JSON as a fenced block: compact and easy to read.

Input: table, list, json, text. Returns: text.
Args:
- `max_cell_chars` (integer; default 120): Truncate long cells to this many characters.
Example: `{"input": [{"plan": "Pro", "price": "$20"}]}` → `"| plan | price |\n|---|---|\n| Pro | $20 |"`

#### to_csv

Render rows as CSV text (columns are the union of all row keys).

Input: table. Returns: text.
Args:
- `delimiter` (string; default ","): Field delimiter.
Example: `{"input": [{"a": 1, "b": "x,y"}]}` → `"a,b\n1,\"x,y\""`

#### to_text

Turn any value into text: JSON is serialized (pretty or compact), text passes through.

Input: json, list, table, number, boolean, text, html. Returns: text.
Args:
- `pretty` (boolean; default false): Indent JSON for readability (costs more tokens).
Example: `{"input": {"a": [1, 2]}}` → `"{\"a\":[1,2]}"`

#### join

Join list items into one text with a separator.

Input: list. Returns: text.
Args:
- `separator` (string; default "\n"): Text placed between items.
Example: `{"input": ["a", "b", "c"], "args": {"separator": ", "}}` → `"a, b, c"`

#### template

Fill a text template like "{{name}} costs {{price}}" from JSON. A table renders once per row, giving a list (or one text with 'join').

Input: json, table, list. Returns: depends on the data.
Args:
- `template` (string; required): Text with {{path}} placeholders: {{name}}, {{address.city}}, {{tags[0]}}. Add a fallback with {{field | n/a}}. In a list of plain values, {{.}} is the item. Inside a recipe, {{names}} that are recipe params are filled by the recipe first.
- `join` (string): For tables and lists: join the rendered rows into one text with this separator (e.g. "\n").
- `on_missing` (string; default "error"; one of error, empty, keep): When a placeholder has no value: error (default, tells you which), empty, or keep the placeholder.
Example: `{"input": {"name": "Ada", "role": {"title": "Engineer"}}, "args": {"template": "{{name}} works as an {{role.title}}."}}` → `"Ada works as an Engineer."`

### compute

#### calc

Evaluate arithmetic exactly: + - * / // % **, comparisons, and functions like round, sqrt, sum, mean, pct. Use this instead of doing math in your head.

Input: no input. Returns: depends on the data.
Args:
- `expr` (string; required): The expression, e.g. "price * qty * (1 + tax/100)". ** for powers (^ is not power).
- `vars` (object): Named values used in expr. Values may be numbers, numeric strings like "$1,200", lists of numbers, or refs to earlier steps ("$2", "$prices[*].amount").
- `places` (integer): Round the result to this many decimal places (half-up).
Example: `{"args": {"expr": "0.1 + 0.2"}}` → `0.3`
Note: Arithmetic is exact (rationals, not floats), so 0.1 + 0.2 is 0.3. Results that aren't whole numbers include the exact decimal in meta.decimal. Functions: abs, round, floor, ceil, sqrt, exp, ln, log, log10, log2, min, max, sum, mean, avg, median, count, pct, pct_change, sin, cos, tan, asin, acos, atan, atan2, pow, hypot, int, trunc, factorial, gcd, lcm, comb, perm. Constants: pi, e, tau, true, false. Integers above 2^53 also come back as meta.exact.

#### now

The current date and time from the server clock, in any time zone: ISO timestamp, weekday, Unix time, week number. Models don't know today's date; this does.

Input: no input. Returns: json.
Args:
- `tz` (string; default "UTC"): Time zone: IANA name ('Asia/Tokyo'), offset ('+05:30'), or common abbreviation ('PST').
Example: `{"args": {"tz": "Europe/Berlin"}}` → `{"iso": "2026-10-02T11:54:17+02:00", "date": "2026-10-02", "weekday": "Friday", "time": "11:54:17", "tz": "Europe/Berlin", "utc_offset": "+02:00", "unix": 1790934857, "day_of_year": 275, "iso_week": 40, "quarter": 4, "i…`

#### date_info

Read a date in almost any format and normalize it: ISO form, weekday, week number, quarter; optionally convert it to another time zone.

Input: no input. Returns: json.
Args:
- `date` (string; required): The date/time: '2026-03-15', 'March 15 2026 3pm', '15/03/2026', a Unix timestamp, or 'now'/'today'.
- `to_tz` (string): Convert to this time zone.
- `assume_tz` (string; default "UTC"): Time zone to assume when the date has none.
- `dayfirst` (boolean; default false): Read 03/04/2026 as 3 April (day first) instead of March 4.
Example: `{"args": {"date": "2026-03-15"}}` → `{"iso": "2026-03-15", "date": "2026-03-15", "weekday": "Sunday", "day_of_year": 74, "iso_week": 11, "quarter": 1, "is_weekend": true, "days_in_month": 31, "is_leap_year": false}`
Note: Ambiguous forms like 03/04/2026 are read month-first unless dayfirst=true, and the result's meta.notes says so. 'now' and 'today' read the server clock.

#### date_add

Add or subtract time from a date: years, months, weeks, days, hours, minutes, seconds, or business days. Month ends are handled (Jan 31 + 1 month = Feb 28/29).

Input: no input. Returns: json.
Args:
- `date` (string; required): Start date/time (any format date_info reads, or 'today'/'now').
- `add` (object; required): Amounts to add; negative to subtract. Keys: years, months, weeks, days, hours, minutes, seconds, business_days. e.g. {"months": 1, "days": -2}.
- `assume_tz` (string; default "UTC"): Time zone to assume when the date has none.
- `dayfirst` (boolean; default false): Read 03/04/2026 as 3 April.
Example: `{"args": {"date": "2026-01-31", "add": {"months": 1}}}` → `{"iso": "2026-02-28", "date": "2026-02-28", "weekday": "Saturday", "day_of_year": 59, "iso_week": 9, "quarter": 1, "is_weekend": true, "days_in_month": 28, "is_leap_year": false}`
Note: Years to days move the calendar (3pm + 1 day is 3pm the next day, even across a DST change); hours, minutes and seconds are elapsed time. Business days are Monday to Friday, counted after the start date like spreadsheet WORKDAY (Saturday + 1 = Monday); public holidays are not considered.

#### date_diff

The time between two dates: total days, weeks, hours, a years/months/days breakdown, and business days.

Input: no input. Returns: json.
Args:
- `start` (string; required): Earlier date/time (any format, or 'today'/'now').
- `end` (string; required): Later date/time. If it's before start, the numbers come out negative.
- `assume_tz` (string; default "UTC"): Time zone to assume when a date has none.
- `dayfirst` (boolean; default false): Read 03/04/2026 as 3 April.
Example: `{"args": {"start": "2026-10-02", "end": "2026-12-25"}}` → `{"days": 84, "weeks": 12, "breakdown": {"years": 0, "months": 2, "days": 23}, "business_days": 60, "direction": "forward"}`
Note: days counts calendar days from start to end (not inclusive of both). business_days counts weekdays after start up to and including end; holidays are not considered.

#### encode

Encode text as base64, base64url, URL (percent) encoding, form encoding, hex, HTML entities, a JSON string literal, \u escapes or rot13.

Input: text, html. Returns: text.
Args:
- `codec` (string; required; one of base64, base64url, url, url_form, hex, html, json_string, unicode_escape, rot13): Which encoding.
Example: `{"input": "hello, world", "args": {"codec": "base64"}}` → `"aGVsbG8sIHdvcmxk"`
Note: Text is encoded as UTF-8 first.

#### decode

Decode base64, base64url, URL (percent) encoding, form encoding, hex, HTML entities, a JSON string literal, \u escapes or rot13 back to text.

Input: text, html. Returns: text.
Args:
- `codec` (string; required; one of base64, base64url, url, url_form, hex, html, json_string, unicode_escape, rot13): Which encoding the input is in.
Example: `{"input": "aGVsbG8sIHdvcmxk", "args": {"codec": "base64"}}` → `"hello, world"`
Note: Base64 padding is optional, and URL-safe characters are accepted for plain base64 too. Bytes that aren't UTF-8 text are refused with a hex preview.

#### url_parse

Split URLs into scheme, host, port, path segments, decoded query parameters and fragment. A list of URLs becomes a table.

Input: text, list. Returns: depends on the data.
Example: `{"input": "https://shop.example.com:8443/items/blue%20mug?id=42&tag=a&tag=b#reviews"}` → `{"url": "https://shop.example.com:8443/items/blue%20mug?id=42&tag=a&tag=b#reviews", "scheme": "https", "host": "shop.example.com", "port": 8443, "path": "/items/blue%20mug", "path_segments": ["items", "blue mug"], "quer…`
Note: A bare domain like example.com/x is read as https. Passwords in URLs are never echoed back.

#### validate

Check data against a JSON Schema and list every violation with its path. Use it to check output before sending it to an API.

Input: json, list, table, text, number, boolean. Returns: json.
Args:
- `schema` (object; required): A JSON Schema (draft 2020-12), e.g. {"type": "object", "required": ["name"]}.
- `max_errors` (integer; default 20): Report at most this many violations.
Example: `{"input": {"name": "Mug", "price": "12"}, "args": {"schema": {"type": "object", "required": ["name", "price", "sku"], "properties": {"name": {"type": "string"}, "price": {"type": "number"}}}}}` → `{"valid": false, "error_count": 2, "errors": [{"path": "$", "message": "'sku' is a required property", "rule": "required"}, {"path": "$.price", "message": "'12' is not of type 'number'", "rule": "type"}]}`
Note: Text input is parsed as JSON first; if it isn't JSON, the result says so instead of failing.

### html

#### html_clean

Clean messy HTML (Word, Google Docs, web pages) at light, medium or aggressive intensity; returns HTML.

Input: html. Returns: html.
Args:
- `level` (string; default "medium"; one of light, medium, aggressive): light: remove metadata, comments and scripts, keep styling. medium: also strip styles, classes, ids, spans and empty tags, tidy tables and lists. aggressive: medium plus unwrap semantic tags (article, section, nav...).
- `drop_boilerplate` (boolean; default false): Remove nav, footers, sidebars, forms, cookie/consent banners and similar before cleaning.
- `images` (string; default "keep"; one of keep, alt, remove): keep: leave <img>. alt: replace each image with [image: alt text]. remove: delete images.
- `normalize_whitespace` (boolean; default true): Collapse runs of whitespace outside <pre>.
Example: `{"input": "<p class=\"MsoNormal\" style=\"margin:0\"><span style=\"font-size:11pt\"><b>Price:</b> $40</span><o:p></o:p></p><!-- c --><p></p>", "args": {"level": "medium"}}` → `"<p><strong>Price:</strong> $40</p>"`
Note: Intensity levels ported from github.com/kevinletchford/html-cleaner (MIT).

#### html_to_text

Turn HTML into readable markdown or plain text, optionally keeping only the main content.

Input: html. Returns: text.
Args:
- `format` (string; default "markdown"; one of markdown, plain): markdown keeps headings, lists, links and tables; plain is bare text with line breaks.
- `main_only` (boolean; default true): Keep only the main content (main/article), dropping navigation, footers, sidebars and banners.
- `keep_links` (boolean; default true): Keep link URLs in markdown output. Turn off to save tokens.
- `keep_images` (boolean; default false): Keep image references in markdown output.
- `base_url` (string): Resolve relative links against this URL. Defaults to the URL the HTML was fetched from.
Example: `{"input": "<html><body><nav><a href='/'>Home</a></nav><main><h1>Title</h1><p>Hello <b>world</b>.</p></main><footer>(c) 2026</footer></body></html>"}` → `"# Title\n\nHello **world**."`

#### select

Pick elements from HTML with a CSS selector; returns their text, HTML or an attribute.

Input: html. Returns: list.
Args:
- `selector` (string; required): A CSS selector, e.g. 'h2', '.price', 'table#rates td', 'a[href*=pdf]'.
- `return` (string; default "text"; one of text, html, attr): text: element text. html: outer HTML. attr: the value of 'attr'.
- `attr` (string): Attribute to read when return='attr', e.g. 'href', 'src', 'content'.
- `limit` (integer; default 100): Maximum number of matches to return.
- `allow_empty` (boolean; default false): Return an empty list instead of an error when nothing matches.
Example: `{"input": "<div><span class='price'>$10</span><span class='price'>$12</span></div>", "args": {"selector": ".price"}}` → `["$10", "$12"]`

#### extract_links

List the links in HTML as rows of {text, href}, with absolute URLs.

Input: html. Returns: table.
Args:
- `contains` (string): Keep only links whose URL or text contains this (case-insensitive).
- `same_domain` (boolean; default false): Keep only links on the same domain as the page.
- `unique` (boolean; default true): Drop duplicate URLs.
- `limit` (integer; default 500): Maximum links to return.
- `base_url` (string): Resolve relative links against this URL. Defaults to the fetched page URL.
Example: `{"input": "<a href='/docs'>Docs</a> <a href='https://x.org/'>X</a> <a href='/docs'>Docs again</a>", "args": {"base_url": "https://example.com/"}}` → `[{"text": "Docs", "href": "https://example.com/docs"}, {"text": "X", "href": "https://x.org/"}]`

#### extract_tables

Turn HTML <table> elements into rows keyed by column header.

Input: html. Returns: depends on the data.
Args:
- `index` (integer): Return only this table (0-based; -1 for last) as a table. Omit to get a list of all tables.
- `min_rows` (integer; default 1): Skip tables with fewer data rows than this (layout tables are often tiny).
Example: `{"input": "<table><tr><th>Plan</th><th>Price</th></tr><tr><td>Free</td><td>$0</td></tr><tr><td>Pro</td><td>$20</td></tr></table>", "args": {"index": 0}}` → `[{"Plan": "Free", "Price": "$0"}, {"Plan": "Pro", "Price": "$20"}]`

#### extract_meta

Get page metadata: title, description, canonical URL, language, Open Graph, and JSON-LD.

Input: html. Returns: json.
Example: `{"input": "<html lang=\"en\"><head><title>Acme</title><meta name=\"description\" content=\"Tools\"><meta property=\"og:image\" content=\"/i.png\"></head><body><h1>Acme Tools</h1></body></html>"}` → `{"title": "Acme", "description": "Tools", "lang": "en", "og": {"image": "/i.png"}, "h1": ["Acme Tools"]}`

### structure

#### outline

Show the heading structure of HTML or markdown, so you can see what's in a document before reading it.

Input: html, text. Returns: table.
Args:
- `max_level` (integer; default 4): Deepest heading level to include (1-6).
Example: `{"input": "# Guide\nintro\n## Install\nsteps\n## Usage\n"}` → `[{"level": 1, "text": "Guide", "offset": 0}, {"level": 2, "text": "Install", "offset": 14}, {"level": 2, "text": "Usage", "offset": 31}]`
Note: For markdown, 'offset' is the character position of the heading, so you can slice from it.

### network

#### fetch

Download a public web page or API response. Returns html, text or json, with source URL, status and time recorded as provenance.

Input: no input. Returns: depends on the data.
Args:
- `url` (string; required): Full http(s) URL.
- `max_bytes` (integer; default 3000000): Stop reading after this many bytes; the result is marked truncated.
- `timeout_s` (number; default 15): Give up after this many seconds.
- `allow_error_status` (boolean; default false): Return 4xx/5xx bodies instead of raising an error.
Example: `{"args": {"url": "https://example.com"}}` → `"(html handle; provenance.final_url, status, content_type, fetched_at)"`
Note: GET only. Private and internal addresses are refused. Binary content (PDF, images) is not supported yet.

### text

#### find

Search text for a word, phrase or regex and return each match with its position and surrounding context.

Input: text, html. Returns: table.
Args:
- `pattern` (string; required): What to look for.
- `regex` (boolean; default false): Treat pattern as a Python regular expression.
- `case_sensitive` (boolean; default false): Match case exactly.
- `context` (integer; default 60): Characters of context to show on each side of a match.
- `limit` (integer; default 20): Maximum matches to return.
Example: `{"input": "Shipping is free over $50.\nReturns accepted within 30 days.", "args": {"pattern": "returns", "context": 10}}` → `[{"offset": 27, "line": 2, "match": "Returns", "context": "over $50. Returns accepted"}]`
Note: Use 'offset' with slice (unit=chars) to read around a match. Finding nothing returns an empty table, not an error.

#### contains

Check whether text (or any item in a list) contains a word, phrase or regex. Returns true or false.

Input: text, html, list. Returns: boolean.
Args:
- `pattern` (string; required): What to look for.
- `regex` (boolean; default false): Treat pattern as a regular expression.
- `case_sensitive` (boolean; default false): Match case exactly.
Example: `{"input": "In stock: yes", "args": {"pattern": "IN STOCK"}}` → `true`

#### slice

Take part of text (by chars, lines, words or approximate tokens) or part of a list (by items). Negative indexes count from the end.

Input: text, html, list. Returns: depends on the data.
Args:
- `start` (integer; default 0): Start index (inclusive).
- `end` (integer): End index (exclusive). Omit to go to the end.
- `unit` (string; default "chars"; one of chars, lines, words, tokens): For text: chars, lines, words or tokens (~4 chars each). Lists always slice by items.
Example: `{"input": "line one\nline two\nline three", "args": {"start": -2, "unit": "lines"}}` → `"line two\nline three"`

#### fit

Shrink text or a list to fit a token budget, marking what was cut. Use before reading anything of unknown size.

Input: text, html, list. Returns: depends on the data.
Args:
- `max_tokens` (integer; required): Token budget (estimated at ~4 chars per token).
- `strategy` (string; default "head"; one of head, tail, head_tail): head: keep the start. tail: keep the end. head_tail: keep both ends and cut the middle.
Example: `{"input": "abcdefghijklmnopqrstuvwxyzabcdefghijklmnopqrstuvwxyzabcdefghijklmnopqrstuvwxyzabcdefghijklmnopqrstuvwxyz", "args": {"max_tokens": 10, "strategy": "head_tail"}}` → `"abcdefghijklmnopqrst\n…[cut 64 chars]…\nghijklmnopqrstuvwxyz"`
Note: Already within budget? The input is returned unchanged. Marker text counts toward the budget only approximately.

#### chunk

Split long text into pieces under a token budget, breaking at paragraphs, lines or sentences where possible.

Input: text, html. Returns: list.
Args:
- `max_tokens` (integer; default 1000): Maximum tokens per chunk (estimated).
- `overlap_tokens` (integer; default 0): Tokens of the previous chunk to repeat at the start of the next.
- `boundary` (string; default "paragraph"; one of paragraph, line, sentence, char): Preferred break point.
Example: `{"input": "Para one is here.\n\nPara two is here.\n\nPara three.", "args": {"max_tokens": 10}}` → `["Para one is here.\n\nPara two is here.", "Para three."]`

#### replace

Replace occurrences of a word, phrase or regex in text.

Input: text, html. Returns: same type as input.
Args:
- `pattern` (string; required): What to replace.
- `replacement` (string; default ""): Replacement text. With regex=true you can use \1 or \g<name>.
- `regex` (boolean; default false): Treat pattern as a regular expression.
- `case_sensitive` (boolean; default true): Match case exactly.
- `count` (integer; default 0): Maximum replacements; 0 means all.
Example: `{"input": "2024-01-05", "args": {"pattern": "(\\d+)-(\\d+)-(\\d+)", "replacement": "\\3/\\2/\\1", "regex": true}}` → `"05/01/2024"`

#### split

Split text into a list by lines, paragraphs, sentences, words or a separator.

Input: text, html. Returns: list.
Args:
- `by` (string; default "lines"; one of lines, paragraphs, sentences, words, separator): How to split.
- `separator` (string): The separator when by='separator'.
- `regex` (boolean; default false): Treat separator as a regular expression.
- `strip` (boolean; default true): Trim whitespace from each piece.
- `drop_empty` (boolean; default true): Drop empty pieces.
Example: `{"input": "a, b,, c", "args": {"by": "separator", "separator": ","}}` → `["a", "b", "c"]`

#### regex_extract

Pull every regex match out of text. One capture group gives a list of strings; several give rows.

Input: text, html. Returns: depends on the data.
Args:
- `pattern` (string; required): Python regular expression.
- `case_sensitive` (boolean; default true): Match case exactly.
- `unique` (boolean; default false): Drop duplicate matches, keeping first-seen order.
- `limit` (integer; default 1000): Maximum matches.
Example: `{"input": "Pro $20/mo, Team $45/mo", "args": {"pattern": "\\$(\\d+)"}}` → `["20", "45"]`
Note: No groups: whole matches. One group: that group. Several groups: rows keyed by group name (or group_1, group_2...).

#### normalize

Tidy text: collapse whitespace, drop blank-line runs, optionally lowercase or Unicode-normalize.

Input: text. Returns: text.
Args:
- `collapse_spaces` (boolean; default true): Turn runs of spaces/tabs into one space.
- `max_blank_lines` (integer; default 1): Allow at most this many consecutive blank lines.
- `lowercase` (boolean; default false): Lowercase everything.
- `unicode_nfkc` (boolean; default false): Apply NFKC normalization (fancy quotes, ligatures, full-width chars).
Example: `{"input": "  Hello\t\tworld  \n\n\n\nBye "}` → `"Hello world\n\nBye"`

#### count

Count chars, words, lines, estimated tokens, items (list entries or object keys), or pattern matches.

Input: text, html, list, json. Returns: number.
Args:
- `unit` (string; default "words"; one of chars, words, lines, tokens_est, items, matches): What to count. Lists default to items.
- `pattern` (string): Pattern to count when unit='matches'.
- `regex` (boolean; default false): Treat pattern as a regular expression.
- `case_sensitive` (boolean; default false): Match case exactly.
Example: `{"input": "the cat and the hat", "args": {"unit": "matches", "pattern": "the"}}` → `2`

#### diff

Compare text against another text; returns counts plus a unified diff.

Input: text, html. Returns: json.
Args:
- `against` (string; required): The other text. In a pipeline this can be a ref like "$0" or a handle id.
- `context` (integer; default 2): Unchanged lines to show around each change.
Example: `{"input": "a\nb\nc", "args": {"against": "a\nB\nc"}}` → `{"identical": false, "added_lines": 1, "removed_lines": 1, "diff": "--- input\n+++ against\n@@ -1,3 +1,3 @@\n a\n-b\n+B\n c"}`
Note: The diff reads from 'input' to 'against': lines starting '+' exist only in 'against'.

### integrity

#### hash

Fingerprint any value (sha256 by default) to check whether content changed or to cite it verbatim.

Input: text, html, json, list, table, number, boolean. Returns: text.
Args:
- `algorithm` (string; default "sha256"; one of sha256, sha1, md5): Hash algorithm.
Example: `{"input": "hello"}` → `"2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824"`
Note: Text is hashed as UTF-8; other types as their compact JSON serialization (the same bytes as a handle's sha256).

## Recipes

Saved pipelines with a public track record. Run one: MCP `run_recipe(name, params)`, HTTP `POST /recipes/{name}/run`.

- **crate_info** (healthy): Facts about a Rust crate from crates.io: latest stable version, description, downloads, repository and last update. Params: `crate`.
- **find_on_page** (healthy): Fetch a page and return each place a phrase appears, with context, without reading the page. Params: `url`, `pattern`.
- **github_readme_outline** (healthy): The heading outline of a GitHub repository's README, to see what it covers before reading it. Params: `repo`, `branch` (default "main").
- **npm_package** (healthy): Facts about a JavaScript package from the npm registry: latest version, license, Node.js requirement, dependencies and repository. Params: `package`.
- **page_tables** (healthy): Fetch a page and return one of its HTML tables as markdown. Params: `url`, `index` (default 0).
- **pypi_package** (healthy): Facts about a Python package from PyPI: latest version, release date, required Python, license, summary and links. Use before suggesting or upgrading a dependency. Params: `package`.
- **page_links** (broken): Fetch a page and list its links, optionally only those containing some text. Params: `url`, `contains` (default "").
- **page_outline** (broken): Fetch a page and return its title, description and heading outline: a cheap look before reading. Params: `url`.
- **page_to_markdown** (broken): Fetch a page and return its main content as markdown, capped to a token budget. Params: `url`, `max_tokens` (default 3000).

## MCP tools

`read_page`, `calculate_expression`, `calculate_date`, `read_package_info` do common jobs directly. `discover_ops`, `run_op`, `run_pipeline`, `read_handle`, `run_recipe`, `save_recipe`, `delete_recipe` reach everything above.

## Errors

Every error is `{"error": {code, detail, suggestion, links}}`; the suggestion names the fix.
