---
title: "AI bot directory: 40 crawlers, tokens and IP lists · DotsAgent"
description: "40 AI crawlers and agents from OpenAI, Anthropic, Google, Apple, Meta, Amazon and others, with robots.txt token, purpose, compliance and IP list."
url: https://dotsagent.io/agent-ready/bots
---

40 crawlers and agents

# AI crawler and agent directory

Each AI bot we could confirm, with the token to use in robots.txt, what it fetches for, whether its operator says it follows the rules, and how to check a request really comes from it. Tokens known only from the community ai.robots.txt registry are marked as such on their own pages.

| Bot | Token | Purpose | Obeys robots.txt | Verify |
| --- | --- | --- | --- | --- |
| [GPTBot](https://dotsagent.io/agent-ready/bots/gptbot)OpenAI | `GPTBot` | Training | — | IP list |
| [OAI-SearchBot](https://dotsagent.io/agent-ready/bots/oai-searchbot)OpenAI | `OAI-SearchBot` | Search | — | IP list |
| [ChatGPT-User](https://dotsagent.io/agent-ready/bots/chatgpt-user)OpenAI | `ChatGPT-User` | User fetches | — | IP list |
| [OAI-AdsBot](https://dotsagent.io/agent-ready/bots/oai-adsbot)OpenAI | `OAI-AdsBot` | Ads | — | IP list |
| [ChatGPT agent](https://dotsagent.io/agent-ready/bots/chatgpt-agent)OpenAI | No token | Agents | — | Signed requests |
| [ClaudeBot](https://dotsagent.io/agent-ready/bots/claudebot)Anthropic | `ClaudeBot` | Training | — | IP list |
| [Claude-SearchBot](https://dotsagent.io/agent-ready/bots/claude-searchbot)Anthropic | `Claude-SearchBot` | Search | — | IP list |
| [Claude-User](https://dotsagent.io/agent-ready/bots/claude-user)Anthropic | `Claude-User` | User fetches | — | IP list |
| [Googlebot](https://dotsagent.io/agent-ready/bots/googlebot)Google | `Googlebot` | Search | — | IP list |
| [Google-Extended](https://dotsagent.io/agent-ready/bots/google-extended)Google | `Google-Extended` | Training | — | — |
| [Google-CloudVertexBot](https://dotsagent.io/agent-ready/bots/google-cloudvertexbot)Google | `Google-CloudVertexBot` | Search | — | IP list |
| [GoogleOther](https://dotsagent.io/agent-ready/bots/googleother)Google | `GoogleOther` | Training | — | IP list |
| [Google-Agent](https://dotsagent.io/agent-ready/bots/google-agent)Google | `Google-Agent` | Agents | — | IP list |
| [Google-GeminiNotebook](https://dotsagent.io/agent-ready/bots/google-gemininotebook)Google | `Google-GeminiNotebook` | User fetches | — | IP list |
| [Gemini Spark](https://dotsagent.io/agent-ready/bots/gemini-spark)Google | No token | Agents | — | — |
| [Gemini-Deep-Research](https://dotsagent.io/agent-ready/bots/gemini-deep-research)Google | `Gemini-Deep-Research` | User fetches | — | — |
| [GoogleAgent-Mariner](https://dotsagent.io/agent-ready/bots/googleagent-mariner)Google | `GoogleAgent-Mariner` | Agents | — | — |
| [Applebot](https://dotsagent.io/agent-ready/bots/applebot)Apple | `Applebot` | Search | — | IP list |
| [Applebot-Extended](https://dotsagent.io/agent-ready/bots/applebot-extended)Apple | `Applebot-Extended` | Training | — | — |
| [PerplexityBot](https://dotsagent.io/agent-ready/bots/perplexitybot)Perplexity | `PerplexityBot` | Search | — | IP list |
| [Perplexity-User](https://dotsagent.io/agent-ready/bots/perplexity-user)Perplexity | `Perplexity-User` | User fetches | — | IP list |
| [meta-externalagent](https://dotsagent.io/agent-ready/bots/meta-externalagent)Meta | `meta-externalagent` | Training | — | — |
| [meta-webindexer](https://dotsagent.io/agent-ready/bots/meta-webindexer)Meta | `meta-webindexer` | Search | — | — |
| [meta-externalfetcher](https://dotsagent.io/agent-ready/bots/meta-externalfetcher)Meta | `meta-externalfetcher` | User fetches | — | — |
| [meta-externalads](https://dotsagent.io/agent-ready/bots/meta-externalads)Meta | `meta-externalads` | Ads | — | — |
| [facebookexternalhit](https://dotsagent.io/agent-ready/bots/facebookexternalhit)Meta | `facebookexternalhit` | User fetches | — | — |
| [Muse](https://dotsagent.io/agent-ready/bots/muse)Meta | No token | Agents | — | — |
| [Amazonbot](https://dotsagent.io/agent-ready/bots/amazonbot)Amazon | `Amazonbot` | Training | — | IP list |
| [Amzn-SearchBot](https://dotsagent.io/agent-ready/bots/amzn-searchbot)Amazon | `Amzn-SearchBot` | Search | — | IP list |
| [Amzn-User](https://dotsagent.io/agent-ready/bots/amzn-user)Amazon | `Amzn-User` | User fetches | — | IP list |
| [DuckAssistBot](https://dotsagent.io/agent-ready/bots/duckassistbot)DuckDuckGo | `DuckAssistBot` | Search | — | IP list |
| [MistralAI-User](https://dotsagent.io/agent-ready/bots/mistralai-user)Mistral | `MistralAI-User` | User fetches | — | IP list |
| [MistralAI-Index](https://dotsagent.io/agent-ready/bots/mistralai-index)Mistral | `MistralAI-Index` | Search | — | IP list |
| [MistralAI-Training](https://dotsagent.io/agent-ready/bots/mistralai-training)Mistral | `MistralAI-Training` | Training | — | — |
| [CCBot](https://dotsagent.io/agent-ready/bots/ccbot)Common Crawl | `CCBot` | Training | — | IP list |
| [cohere-ai](https://dotsagent.io/agent-ready/bots/cohere-ai)Cohere | `cohere-ai` | User fetches | — | — |
| [Bytespider](https://dotsagent.io/agent-ready/bots/bytespider)ByteDance | `Bytespider` | Training | — | — |
| [YouBot](https://dotsagent.io/agent-ready/bots/youbot)You.com | `YouBot` | Search | — | — |
| [Diffbot](https://dotsagent.io/agent-ready/bots/diffbot)Diffbot | `Diffbot` | Search | — | — |
| [Timpibot](https://dotsagent.io/agent-ready/bots/timpibot)Timpi | `Timpibot` | Training | — | — |

## What the purposes mean

- **Training**: Collects pages for model training or general research crawling. Google-Extended and Applebot-Extended are control tokens that never fetch anything themselves.
- **Search**: Indexes pages for a search product, from Google Search to AI answer engines such as ChatGPT search, Perplexity and DuckAssist. Blocking one keeps you out of that product.
- **User fetches**: Fetches a page because a person asked for it in a chat or tool. Several operators say robots.txt may not apply to these requests.
- **Agents**: Browses and acts on a user's behalf, often in a full browser. Many have no token and can be identified only by signature, if at all.
- **Ads**: Checks ad landing pages that advertisers submitted. These are left out of the robots.txt generator.

## Check a bot is who it says it is

A user-agent string costs nothing to fake. Most large operators publish the IP ranges their bots use as JSON files, marked IP list in the table. Compare the request's address with the list for that token, and treat a match on the name alone as unverified.

Some operators also document reverse DNS. Common Crawl, for example, says CCBot hosts resolve to *.crawl.commoncrawl.org; confirm that the hostname resolves back to the same address before you trust it.

Agents increasingly sign requests with HTTP Message Signatures (RFC 9421) under the Web Bot Auth draft. The request names its operator in Signature-Agent, and you verify it with the public keys at that domain's /.well-known/http-message-signatures-directory. A valid signature proves the operator, not the user behind it or what that user may do.

[Build a robots.txt from this list →](https://dotsagent.io/agent-ready/robots-txt)

## Questions about AI crawlers

### Which AI crawlers ignore robots.txt?

By their operators' own documentation, ChatGPT-User, Perplexity-User, Google-Agent, Google-GeminiNotebook, Amzn-User, meta-externalfetcher and facebookexternalhit may not follow it, mostly because a user triggered the fetch. Bytespider is reported by the community not to respect it.

### What is the difference between GPTBot, OAI-SearchBot and ChatGPT-User?

GPTBot collects pages for training OpenAI models, OAI-SearchBot indexes pages for ChatGPT search, and ChatGPT-User fetches a page when a person asks for it. The first two follow robots.txt; OpenAI says its rules may not apply to the third.

### Which AI agents have no robots.txt token?

ChatGPT agent, Gemini Spark and Meta's Muse. ChatGPT agent can be verified through its Web Bot Auth signature from https://chatgpt.com. Gemini Spark in local Chrome mode looks like the user's own browser, and Muse has no published user agent.

### How do I verify a request really comes from Googlebot?

Google publishes Googlebot's IP ranges as a JSON file, linked from its entry in this directory. Check the request's address against those ranges rather than trusting the user-agent string.

### Where does this list of AI bots come from?

Each entry links to the operator's documentation where it exists, and tokens without it come from the community ai.robots.txt registry. We left out DeepSeekBot, Kimi-User, Manus-User, NovaAct, Devin and PetalBot because we found no operator documentation for them. The list was checked on 1 October 2026.

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026
