Make your site readable by AI agents
AI crawlers and agents read your site differently from people. The tools below generate the files they look for, and the reference after them says which conventions are standards, which are drafts and how many clients actually use them.
llms.txt generator
Enter a name, a summary and sections of links, and copy a file in the llmstxt.org format.
robots.txt for AI bots
Set a policy per bot purpose, flip single crawlers, add Content-Signal and download the result.
API catalog
List your APIs with their OpenAPI files and docs, and get an RFC 9727 linkset plus an nginx block to serve it.
AI bot directory
40 crawlers and agents with their tokens, purposes, robots.txt behaviour and published IP lists.
The conventions, and how settled each one is
Status as of October 2026. Each entry links to the spec or to the evidence behind it, so you can judge what is worth the effort on your site.
robots.txt Standard
The file at your site root names crawlers by token and lists the paths each one should skip; it is standardised as RFC 9309. AI operators now publish separate tokens for training, search and user-triggered fetches, so a single file can refuse training and still welcome search. Compliance is voluntary, and several user-triggered fetchers say the rules may not apply to them.
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /Content-Signal Widely used
One extra robots.txt line that states how fetched content may be used: search, ai-input and ai-train, each set to yes or no. Cloudflare released it as a CC0 policy in September 2025 and serves it on about 3.8 million domains through its managed robots.txt. The matching IETF draft expired on 4 April 2026, while the IETF AIPREF working group carries on with a vocabulary draft.
User-agent: *
Content-Signal: ai-train=no, search=yes, ai-input=yes
Allow: /llms.txt Proposal
A Markdown file at /llms.txt with an H1 name, a blockquote summary and H2 sections of links, so an agent can reach your key pages without parsing navigation. The spec at llmstxt.org was revised on 10 August 2026, and llms-full.txt is not part of it. We have no operator statement that any major crawler reads the file.
# Acme API
> Card payments for small online shops.
## Docs
- [Quickstart](https://example.com/docs/quickstart.md): first requestMarkdown versions of pages Partial adoption
Publish a Markdown copy of each page at page.md or page.html.md, or serve Markdown to any request whose Accept header asks for text/markdown. In Checkly's February 2026 test, Claude Code, Cursor and OpenCode sent that header, while Codex, Gemini CLI, GitHub Copilot and Windsurf did not. AnswerShare counted 1,277,965 requests from frontier crawlers and none of them asked for Markdown.
GET /docs/quickstart HTTP/1.1
Host: example.com
Accept: text/markdownAPI catalog Standard
RFC 9727, published in June 2025, reserves /.well-known/api-catalog for a list of your public APIs, each linked to its machine-readable description, docs and status page. The response must be application/linkset+json, and other pages can point to it with the api-catalog link relation.
{
"linkset": [
{
"anchor": "https://api.example.com/v1",
"service-desc": [{ "href": "https://api.example.com/v1/openapi.json" }]
}
]
}Signed agent requests (Web Bot Auth) Draft
Agents sign each request with HTTP Message Signatures (RFC 9421) and name their operator in a Signature-Agent header. You check the signature against the keys that operator publishes under /.well-known/http-message-signatures-directory. The IETF working group draft dates from 1 September 2026; ChatGPT agent already signs as https://chatgpt.com, and Google is experimenting with https://agent.bot.goog.
Signature-Agent: "https://chatgpt.com"NLWeb Stalled
Microsoft's NLWeb gives a site /ask and /mcp endpoints that agents can query in natural language, and its spec is at version 0.55. Lumar reports that the nlweb.ai certificate expired on 4 August 2026 and that the repository has been quiet since June, so wait before building on it.
Agents that don't announce themselves
Some agents publish no robots.txt token and browse as the user would, so no User-agent rule can reach them. ChatGPT agent signs its requests with Web Bot Auth and sends Signature-Agent: "https://chatgpt.com", which you can check against OpenAI's key directory; Cloudflare labels that traffic chatgpt-agent. Gemini Spark in local Chrome mode and Meta's Muse have no documented signal at all.
Questions about agent-ready sites
What makes a website agent-ready?
No single standard defines it. In practice it means a robots.txt that names AI bots by purpose, optional files such as llms.txt and an API catalog, Markdown versions of key pages, and a way to verify the agents you let in. Of these, only robots.txt and the API catalog are published RFCs.
Do AI crawlers read llms.txt?
We have no operator documentation saying that any crawler in our directory fetches it. Lighthouse checks whether the file exists, and an agent you point at it can read it like any other page. Treat it as a convenience for agents, not a search feature.
Should I block AI crawlers on my site?
Decide per purpose rather than all at once. Training tokens such as GPTBot, ClaudeBot and Google-Extended can be refused without touching search. Blocking search tokens such as OAI-SearchBot or PerplexityBot keeps you out of those products entirely.
Can I block ChatGPT agent with robots.txt?
No, because it publishes no token. You can recognise it by its Web Bot Auth signature from https://chatgpt.com and allow or deny it at your server or CDN. A valid signature proves which operator sent the request, not who the user is or what they may do.
Is Content-Signal an official standard?
No. It is a CC0 policy from Cloudflare, launched in September 2025 and served on about 3.8 million domains through Cloudflare's managed robots.txt. Its IETF draft expired on 4 April 2026, and the AIPREF working group is still drafting a shared vocabulary.