dotsagent.io
Language:English
Open your site · generator

robots.txt generator for AI bots

Pick a policy for each kind of AI bot, override single crawlers and copy a file that follows RFC 9309. Training, search and user-triggered fetchers use separate tokens, so you can refuse one and welcome another.

Start from
Training

Blocking these opts your pages out of model training by these operators and leaves your search listings alone.

Search

Blocking these also blocks Googlebot and Applebot, which takes you out of Google Search and Apple's Siri, Spotlight and Safari results.

User fetches

These fetch a page because a person asked for it, and several operators say robots.txt may not apply, so a block here is a request rather than a guarantee.

Agents

Google-Agent generally ignores robots.txt and most browsing agents have no token at all, so this group changes little in practice.

Tap a bot to flip it against its group's setting.

Content-Signal
search
ai-input
ai-train

Content-Signal says how fetched pages may be used: search for search results, ai-input for feeding AI answers and ai-train for model training. It comes from Cloudflare's CC0 policy and states a preference without blocking anything.

/robots.txt
# AI crawlers blocked from the whole site
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: GoogleOther
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Amazonbot
User-agent: MistralAI-Training
User-agent: CCBot
User-agent: Bytespider
User-agent: Timpibot
Disallow: /

# Allowed AI crawlers, private paths excluded
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: Googlebot
User-agent: Google-CloudVertexBot
User-agent: Google-Agent
User-agent: Google-GeminiNotebook
User-agent: Gemini-Deep-Research
User-agent: GoogleAgent-Mariner
User-agent: Applebot
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: meta-webindexer
User-agent: meta-externalfetcher
User-agent: facebookexternalhit
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: DuckAssistBot
User-agent: MistralAI-User
User-agent: MistralAI-Index
User-agent: cohere-ai
User-agent: YouBot
User-agent: Diffbot
Disallow: /admin/
Disallow: /api/

# All other crawlers
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Disallow: /admin/
Disallow: /api/

Sitemap: https://example.com/sitemap.xml

robots.txt is voluntary. ChatGPT-User, Perplexity-User and other user-triggered fetchers say the rules may not apply to them, and some agents have no token at all. To enforce a block, match requests against published IP lists or signatures and deny them at your firewall or WAF.

See every bot in the directory →

How the generated file works

A crawler follows its own group

A crawler that finds its token in a User-agent line follows that group and skips the rules under User-agent: *. That is why the generator repeats your private paths for the bots you allow by name, and gives blocked bots a single Disallow: /.

Some tokens never crawl

Google-Extended and Applebot-Extended have no user agent of their own. Googlebot and Applebot do the fetching, and the extended tokens tell Google and Apple whether that content may train their models; Google also applies Google-Extended to Gemini grounding.

Training and search are split by most operators

OpenAI, Anthropic, Amazon and Mistral each run separate tokens for training and for search, so the No training preset keeps you in their search products. Meta is the exception: meta-externalagent covers both training and indexing.

Applebot falls back to Googlebot

If your file has no Applebot group, Apple says Applebot follows your Googlebot rules instead. Applebot also ignores Crawl-delay, as does Amazonbot, while Anthropic says ClaudeBot honours it.

Questions about robots.txt and AI bots

How do I block GPTBot but stay in ChatGPT search?

Disallow GPTBot and leave OAI-SearchBot allowed. GPTBot collects training data, OAI-SearchBot indexes pages for ChatGPT search, and ChatGPT-User handles fetches a person asks for. OpenAI documents all three on one page, each with its own IP list.

Does blocking Google-Extended remove my site from Google Search?

Google-Extended is a control token with no user agent of its own; Googlebot does the crawling. Disallowing it tells Google not to use your content for Gemini training and grounding, while your Googlebot rules decide what is crawled for Search.

Do AI crawlers respect Crawl-delay?

Some do. Anthropic says ClaudeBot honours it, while Apple and Amazon say Applebot and Amazonbot do not. Rate limiting at your server is the dependable option.

Why does ChatGPT-User still visit after I blocked it?

OpenAI says robots.txt rules may not apply to ChatGPT-User, because a person asked for the page. Perplexity-User, Amzn-User and meta-externalfetcher come with similar caveats. Where the operator publishes an IP list, match requests against it and block at your firewall.

What does Content-Signal do in robots.txt?

It states how content may be used once fetched, through three signals: search, ai-input and ai-train, each set to yes or no. It expresses a preference and blocks nothing. Cloudflare's Markdown for Agents adds ai-train=yes, search=yes, ai-input=yes by default, so check its output if you use that feature.

Sources

  1. rfc-editor.org/rfc/rfc9309
  2. contentsignals.org/
  3. developers.google.com/search/docs/crawling-indexing/google-common-crawlers

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026