dotsagent.io
Language:English
Open your site · reference

Make your site readable by AI agents

AI crawlers and agents read your site differently from people. The tools below generate the files they look for, and the reference after them says which conventions are standards, which are drafts and how many clients actually use them.

The conventions, and how settled each one is

Status as of October 2026. Each entry links to the spec or to the evidence behind it, so you can judge what is worth the effort on your site.

robots.txt Standard

The file at your site root names crawlers by token and lists the paths each one should skip; it is standardised as RFC 9309. AI operators now publish separate tokens for training, search and user-triggered fetches, so a single file can refuse training and still welcome search. Compliance is voluntary, and several user-triggered fetchers say the rules may not apply to them.

/robots.txt
User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

platform.openai.comdevelopers.google.com

Content-Signal Widely used

One extra robots.txt line that states how fetched content may be used: search, ai-input and ai-train, each set to yes or no. Cloudflare released it as a CC0 policy in September 2025 and serves it on about 3.8 million domains through its managed robots.txt. The matching IETF draft expired on 4 April 2026, while the IETF AIPREF working group carries on with a vocabulary draft.

/robots.txt
User-agent: *
Content-Signal: ai-train=no, search=yes, ai-input=yes
Allow: /

blog.cloudflare.comdatatracker.ietf.org

llms.txt Proposal

A Markdown file at /llms.txt with an H1 name, a blockquote summary and H2 sections of links, so an agent can reach your key pages without parsing navigation. The spec at llmstxt.org was revised on 10 August 2026, and llms-full.txt is not part of it. We have no operator statement that any major crawler reads the file.

/llms.txt
# Acme API

> Card payments for small online shops.

## Docs

- [Quickstart](https://example.com/docs/quickstart.md): first request

llmstxt.orgllmstxt.org

Markdown versions of pages Partial adoption

Publish a Markdown copy of each page at page.md or page.html.md, or serve Markdown to any request whose Accept header asks for text/markdown. In Checkly's February 2026 test, Claude Code, Cursor and OpenCode sent that header, while Codex, Gemini CLI, GitHub Copilot and Windsurf did not. AnswerShare counted 1,277,965 requests from frontier crawlers and none of them asked for Markdown.

Request
GET /docs/quickstart HTTP/1.1
Host: example.com
Accept: text/markdown

developers.cloudflare.comchecklyhq.comanswershare.com

API catalog Standard

RFC 9727, published in June 2025, reserves /.well-known/api-catalog for a list of your public APIs, each linked to its machine-readable description, docs and status page. The response must be application/linkset+json, and other pages can point to it with the api-catalog link relation.

/.well-known/api-catalog
{
  "linkset": [
    {
      "anchor": "https://api.example.com/v1",
      "service-desc": [{ "href": "https://api.example.com/v1/openapi.json" }]
    }
  ]
}

rfc-editor.org

Signed agent requests (Web Bot Auth) Draft

Agents sign each request with HTTP Message Signatures (RFC 9421) and name their operator in a Signature-Agent header. You check the signature against the keys that operator publishes under /.well-known/http-message-signatures-directory. The IETF working group draft dates from 1 September 2026; ChatGPT agent already signs as https://chatgpt.com, and Google is experimenting with https://agent.bot.goog.

Request header
Signature-Agent: "https://chatgpt.com"

datatracker.ietf.orgblog.cloudflare.com

NLWeb Stalled

Microsoft's NLWeb gives a site /ask and /mcp endpoints that agents can query in natural language, and its spec is at version 0.55. Lumar reports that the nlweb.ai certificate expired on 4 August 2026 and that the repository has been quiet since June, so wait before building on it.

github.comagentic-readiness.lumar.io

Agents that don't announce themselves

Some agents publish no robots.txt token and browse as the user would, so no User-agent rule can reach them. ChatGPT agent signs its requests with Web Bot Auth and sends Signature-Agent: "https://chatgpt.com", which you can check against OpenAI's key directory; Cloudflare labels that traffic chatgpt-agent. Gemini Spark in local Chrome mode and Meta's Muse have no documented signal at all.

Questions about agent-ready sites

What makes a website agent-ready?

No single standard defines it. In practice it means a robots.txt that names AI bots by purpose, optional files such as llms.txt and an API catalog, Markdown versions of key pages, and a way to verify the agents you let in. Of these, only robots.txt and the API catalog are published RFCs.

Do AI crawlers read llms.txt?

We have no operator documentation saying that any crawler in our directory fetches it. Lighthouse checks whether the file exists, and an agent you point at it can read it like any other page. Treat it as a convenience for agents, not a search feature.

Should I block AI crawlers on my site?

Decide per purpose rather than all at once. Training tokens such as GPTBot, ClaudeBot and Google-Extended can be refused without touching search. Blocking search tokens such as OAI-SearchBot or PerplexityBot keeps you out of those products entirely.

Can I block ChatGPT agent with robots.txt?

No, because it publishes no token. You can recognise it by its Web Bot Auth signature from https://chatgpt.com and allow or deny it at your server or CDN. A valid signature proves which operator sent the request, not who the user is or what they may do.

Is Content-Signal an official standard?

No. It is a CC0 policy from Cloudflare, launched in September 2025 and served on about 3.8 million domains through Cloudflare's managed robots.txt. Its IETF draft expired on 4 April 2026, and the AIPREF working group is still drafting a shared vocabulary.

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026