dotsagent.io
Language:English
40 crawlers and agents

AI crawler and agent directory

Each AI bot we could confirm, with the token to use in robots.txt, what it fetches for, whether its operator says it follows the rules, and how to check a request really comes from it. Tokens known only from the community ai.robots.txt registry are marked as such on their own pages.

BotTokenPurposeObeys robots.txtVerify
GPTBot
OpenAI
GPTBotTrainingYesIP list
OAI-SearchBot
OpenAI
OAI-SearchBotSearchYesIP list
ChatGPT-User
OpenAI
ChatGPT-UserUser fetchesNoIP list
OAI-AdsBot
OpenAI
OAI-AdsBotAdsNot statedIP list
ChatGPT agent
OpenAI
No tokenAgentsNot statedSigned requests
ClaudeBot
Anthropic
ClaudeBotTrainingYesIP list
Claude-SearchBot
Anthropic
Claude-SearchBotSearchYesIP list
Claude-User
Anthropic
Claude-UserUser fetchesYesIP list
Googlebot
Google
GooglebotSearchYesIP list
Google-Extended
Google
Google-ExtendedTrainingYes—
Google-CloudVertexBot
Google
Google-CloudVertexBotSearchYesIP list
GoogleOther
Google
GoogleOtherTrainingYesIP list
Google-Agent
Google
Google-AgentAgentsNoIP list
Google-GeminiNotebook
Google
Google-GeminiNotebookUser fetchesNoIP list
Gemini Spark
Google
No tokenAgentsNot stated—
Gemini-Deep-Research
Google
Gemini-Deep-ResearchUser fetchesNot stated—
GoogleAgent-Mariner
Google
GoogleAgent-MarinerAgentsNot stated—
Applebot
Apple
ApplebotSearchYesIP list
Applebot-Extended
Apple
Applebot-ExtendedTrainingYes—
PerplexityBot
Perplexity
PerplexityBotSearchYesIP list
Perplexity-User
Perplexity
Perplexity-UserUser fetchesNoIP list
meta-externalagent
Meta
meta-externalagentTrainingYes—
meta-webindexer
Meta
meta-webindexerSearchYes—
meta-externalfetcher
Meta
meta-externalfetcherUser fetchesNo—
meta-externalads
Meta
meta-externaladsAdsYes—
facebookexternalhit
Meta
facebookexternalhitUser fetchesNo—
Muse
Meta
No tokenAgentsNot stated—
Amazonbot
Amazon
AmazonbotTrainingYesIP list
Amzn-SearchBot
Amazon
Amzn-SearchBotSearchYesIP list
Amzn-User
Amazon
Amzn-UserUser fetchesNoIP list
DuckAssistBot
DuckDuckGo
DuckAssistBotSearchYesIP list
MistralAI-User
Mistral
MistralAI-UserUser fetchesYesIP list
MistralAI-Index
Mistral
MistralAI-IndexSearchYesIP list
MistralAI-Training
Mistral
MistralAI-TrainingTrainingYes—
CCBot
Common Crawl
CCBotTrainingYesIP list
cohere-ai
Cohere
cohere-aiUser fetchesNot stated—
Bytespider
ByteDance
BytespiderTrainingNo—
YouBot
You.com
YouBotSearchNot stated—
Diffbot
Diffbot
DiffbotSearchNot stated—
Timpibot
Timpi
TimpibotTrainingNot stated—

What the purposes mean

Training
Collects pages for model training or general research crawling. Google-Extended and Applebot-Extended are control tokens that never fetch anything themselves.
Search
Indexes pages for a search product, from Google Search to AI answer engines such as ChatGPT search, Perplexity and DuckAssist. Blocking one keeps you out of that product.
User fetches
Fetches a page because a person asked for it in a chat or tool. Several operators say robots.txt may not apply to these requests.
Agents
Browses and acts on a user's behalf, often in a full browser. Many have no token and can be identified only by signature, if at all.
Ads
Checks ad landing pages that advertisers submitted. These are left out of the robots.txt generator.

Check a bot is who it says it is

A user-agent string costs nothing to fake. Most large operators publish the IP ranges their bots use as JSON files, marked IP list in the table. Compare the request's address with the list for that token, and treat a match on the name alone as unverified.

Some operators also document reverse DNS. Common Crawl, for example, says CCBot hosts resolve to *.crawl.commoncrawl.org; confirm that the hostname resolves back to the same address before you trust it.

Agents increasingly sign requests with HTTP Message Signatures (RFC 9421) under the Web Bot Auth draft. The request names its operator in Signature-Agent, and you verify it with the public keys at that domain's /.well-known/http-message-signatures-directory. A valid signature proves the operator, not the user behind it or what that user may do.

Build a robots.txt from this list →

Questions about AI crawlers

Which AI crawlers ignore robots.txt?

By their operators' own documentation, ChatGPT-User, Perplexity-User, Google-Agent, Google-GeminiNotebook, Amzn-User, meta-externalfetcher and facebookexternalhit may not follow it, mostly because a user triggered the fetch. Bytespider is reported by the community not to respect it.

What is the difference between GPTBot, OAI-SearchBot and ChatGPT-User?

GPTBot collects pages for training OpenAI models, OAI-SearchBot indexes pages for ChatGPT search, and ChatGPT-User fetches a page when a person asks for it. The first two follow robots.txt; OpenAI says its rules may not apply to the third.

Which AI agents have no robots.txt token?

ChatGPT agent, Gemini Spark and Meta's Muse. ChatGPT agent can be verified through its Web Bot Auth signature from https://chatgpt.com. Gemini Spark in local Chrome mode looks like the user's own browser, and Muse has no published user agent.

How do I verify a request really comes from Googlebot?

Google publishes Googlebot's IP ranges as a JSON file, linked from its entry in this directory. Check the request's address against those ranges rather than trusting the user-agent string.

Where does this list of AI bots come from?

Each entry links to the operator's documentation where it exists, and tokens without it come from the community ai.robots.txt registry. We left out DeepSeekBot, Kimi-User, Manus-User, NovaAct, Devin and PetalBot because we found no operator documentation for them. The list was checked on 1 October 2026.

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026