> Documentation index: https://unbrowse.ai/llms.txt. Fetch it to find every page.

# Make your site agent-ready

Agents reach your site through Unbrowse as tools: "search products", "get an order", "list events". Unbrowse reads what your site already publishes for machines first, and only records a browser when you publish nothing. The more you declare, the more complete, accurate and fast those tools are, and the fewer wrong guesses anyone makes about your site.

Check your site on its page, `https://unbrowse.ai/sites/<your-domain>` ("Check now"), or with `GET https://unbrowse.ai/api/v1/sites/<your-domain>/readiness?refresh=1`. It returns a score out of 100, what was found, and the fix for each gap. A refresh runs at most once every 10 minutes per site, and reads only public URLs.

## What Unbrowse reads, in order

| Signal | Points | What it gives agents |
|---|---|---|
| Your own MCP server | 25 open, 20 with OAuth | Your tools, under your names. Unbrowse proxies the read-only tools of an open server (`readOnlyHint: true`) and never a destructive one. |
| An API description: OpenAPI, or GraphQL with introspection | 20 | One tool per public GET operation, or per root query field, with your names, parameters and descriptions. |
| An OpenSearch description | 10 | A proven search tool in one request, no browser. |
| A sitemap listed in robots.txt | 10 | Every page type, templated into a reader tool. |
| Structured data (JSON-LD, or JSON in the HTML) | 10 | Clean records instead of scraped text. |
| Content in the HTML | 10 | Plain HTTP reads (about 0.1–0.5 s), not a rendered browser (seconds). |
| No bot wall on public reads | 10 | Tools that keep working. |
| `/llms.txt` | 5 | The pages and docs agents should read. |
| `/.well-known/unbrowse.json` | 5 | Your own pointers, and the paths to skip. |

## Declare your API and MCP server

Publish your OpenAPI document at `/openapi.json`, or point to it:

```json
// https://example.com/.well-known/unbrowse.json
{
  "version": 1,
  "openapi": "/api/openapi.json",
  "graphql": "/api/graphql",
  "mcp": "https://mcp.example.com/mcp",
  "sitemaps": ["/sitemap-products.xml"],
  "exclude": ["/api/nav", "/wp-json/menus"]
}
```

- **openapi**: Unbrowse turns each GET operation with no security requirement into a tool. A required parameter needs an `example`, an `enum` or a `default`, because each tool is proven with one real call before it is listed. Operations that need a key, a cookie or a header, and every write, are never listed.
- **graphql**: a GraphQL endpoint (Unbrowse also tries `/graphql`, `/api/graphql` and `https://api.<your-domain>/graphql`). With introspection on, each root query field becomes a tool: its scalar arguments become inputs, and a selection of the returned type's fields (two levels, following `edges`, `nodes` and lists) is asked for. A required argument needs a default in the schema; page sizes (`first`, `limit`…) get 10. Fields that depend on a session (`viewer`, `me`) and Relay's `node` are skipped. Servers that cap query depth are read with smaller `__type` lookups.
- **mcp**: a streamable-HTTP MCP server. Unbrowse also looks at `/mcp`, `https://mcp.<your-domain>/mcp` and `/.well-known/mcp/server-card.json`. An open server's read-only tools are listed and called through Unbrowse. A server behind OAuth is recorded and shown as your official server; people sign in to it directly.
- **exclude**: path prefixes that are page chrome (menus, widgets, tracking). A recorded request there is never turned into a tool.

## Put state in the URL

- Search, filters, sort and paging as query parameters: `/search?q=dune&type=movie&page=2`. A tool's inputs are the parts of the request that change. State kept only in the page's memory, or sent as a POST body built from the UI, produces tools pinned to one recording, which Unbrowse refuses to publish.
- Detail pages at stable, predictable paths: `/products/{id}`, not `/p?ref=8f3a…`.
- Link each result to its detail page.

## Use real forms and labels

A `<form method="get">` with named inputs (`name="q"`, `name="location"`) and `<label>`s becomes a tool with no browser, named after your fields. An autocomplete that only emits a GraphQL blob does not.

## Keep the API path obvious

- One request per user action, returning JSON the page renders.
- Separate page chrome from content: header, navigation, widgets and tracking on their own paths, or listed in `exclude`.
- CSRF tokens and signed parameters are fine if a public endpoint hands them out. One-time values baked into the HTML break replay.
- No build hashes in data URLs (`/_next/data/<buildId>/…`).
- Answer failures with honest status codes (403, 404, JSON errors). Unbrowse takes a tool down after two failures your site answers, so good tools stay and broken ones leave quickly.

## Let verified agents through

Keep public read paths open to agents, or allow Unbrowse's indexer through your bot manager. Sign-in should be a normal form (username, password, one-time code), so a person's saved login can fill it and their session carries the actions behind it. Unbrowse never publishes anything behind a sign-in.

## What happens next

Once your site declares these, the next index pass (or `unbrowse.index` / `POST /api/v1/index`) publishes your tools, proven. Each is re-checked on a schedule with the input your document gives, and taken down if your site stops answering. Your site then has its own MCP server at `https://unbrowse.ai/mcp/<your-domain>` and a skill anyone can install: `npx skills add https://unbrowse.ai --skill <your-site>`.
