AI 爬虫和代理目录
这里列出了我们能够确认的所有 AI 机器人,包括 robots.txt 中使用的令牌、抓取目的、运营方是否声明遵守相关规则,以及如何验证请求确实来自该机器人。令牌仅见于社区维护的 ai.robots.txt 登记表时,会在对应页面特别标注。
| 机器人 | 令牌 | 用途 | 遵守 robots.txt | 验证 |
|---|---|---|---|---|
| GPTBot | GPTBot | 训练 | 是 | IP 列表 |
| OAI-SearchBot | OAI-SearchBot | 搜索 | 是 | IP 列表 |
| ChatGPT-User | ChatGPT-User | 用户请求 | 否 | IP 列表 |
| OAI-AdsBot | OAI-AdsBot | 广告 | 未说明 | IP 列表 |
| ChatGPT agent | 无令牌 | Agent | 未说明 | 签名请求 |
| ClaudeBot | ClaudeBot | 训练 | 是 | IP 列表 |
| Claude-SearchBot | Claude-SearchBot | 搜索 | 是 | IP 列表 |
| Claude-User | Claude-User | 用户请求 | 是 | IP 列表 |
| Googlebot | Googlebot | 搜索 | 是 | IP 列表 |
| Google-Extended | Google-Extended | 训练 | 是 | — |
| Google-CloudVertexBot | Google-CloudVertexBot | 搜索 | 是 | IP 列表 |
| GoogleOther | GoogleOther | 训练 | 是 | IP 列表 |
| Google-Agent | Google-Agent | Agent | 否 | IP 列表 |
| Google-GeminiNotebook | Google-GeminiNotebook | 用户请求 | 否 | IP 列表 |
| Gemini Spark | 无令牌 | Agent | 未说明 | — |
| Gemini-Deep-Research | Gemini-Deep-Research | 用户请求 | 未说明 | — |
| GoogleAgent-Mariner | GoogleAgent-Mariner | Agent | 未说明 | — |
| Applebot | Applebot | 搜索 | 是 | IP 列表 |
| Applebot-Extended | Applebot-Extended | 训练 | 是 | — |
| PerplexityBot | PerplexityBot | 搜索 | 是 | IP 列表 |
| Perplexity-User | Perplexity-User | 用户请求 | 否 | IP 列表 |
| meta-externalagent | meta-externalagent | 训练 | 是 | — |
| meta-webindexer | meta-webindexer | 搜索 | 是 | — |
| meta-externalfetcher | meta-externalfetcher | 用户请求 | 否 | — |
| meta-externalads | meta-externalads | 广告 | 是 | — |
| facebookexternalhit | facebookexternalhit | 用户请求 | 否 | — |
| Muse | 无令牌 | Agent | 未说明 | — |
| Amazonbot | Amazonbot | 训练 | 是 | IP 列表 |
| Amzn-SearchBot | Amzn-SearchBot | 搜索 | 是 | IP 列表 |
| Amzn-User | Amzn-User | 用户请求 | 否 | IP 列表 |
| DuckAssistBot | DuckAssistBot | 搜索 | 是 | IP 列表 |
| MistralAI-User | MistralAI-User | 用户请求 | 是 | IP 列表 |
| MistralAI-Index | MistralAI-Index | 搜索 | 是 | IP 列表 |
| MistralAI-Training | MistralAI-Training | 训练 | 是 | — |
| CCBot | CCBot | 训练 | 是 | IP 列表 |
| cohere-ai | cohere-ai | 用户请求 | 未说明 | — |
| Bytespider | Bytespider | 训练 | 否 | — |
| YouBot | YouBot | 搜索 | 未说明 | — |
| Diffbot | Diffbot | 搜索 | 未说明 | — |
| Timpibot | Timpibot | 训练 | 未说明 | — |
用途说明
- 训练
- 收集网页,用于模型训练或一般研究爬取。Google-Extended 和 Applebot-Extended 是控制令牌,本身不会抓取任何内容。
- 搜索
- 为搜索产品编入网页索引,涵盖 Google Search 以及 ChatGPT search、Perplexity 和 DuckAssist 等 AI 问答引擎。屏蔽此类机器人,就无法出现在对应产品中。
- 用户请求
- 根据用户在聊天或工具中的请求抓取网页。多家运营方表示,robots.txt 可能不适用于这类请求。
- Agent
- 代表用户浏览网页并执行操作,通常使用完整浏览器。许多代理没有令牌,因此只能通过签名识别(如果能够识别的话)。
- 广告
- 检查广告主提交的广告落地页。这些机器人不会包含在 robots.txt 生成器中。
验证爬虫身份
伪造 User-Agent 字符串毫不费力。大多数大型运营方会以 JSON 文件公布其爬虫使用的 IP 地址段,表格中标记为「IP 列表」。将请求的 IP 地址与该 token 对应的列表进行比对;仅凭名称匹配不能视为已验证。
部分运营方还会说明反向 DNS 信息。例如,Common Crawl 表示 CCBot 主机名会解析到 *.crawl.commoncrawl.org;在信任该主机前,请确认该主机名解析回同一个 IP 地址。
越来越多的智能体依据 Web Bot Auth 草案,使用 HTTP Message Signatures (RFC 9421) 对请求进行签名。请求通过 Signature-Agent 声明其运营方,你可以使用该域名下 /.well-known/http-message-signatures-directory 中的公钥验证签名。有效签名只能证明运营方身份,无法证明背后的用户是谁,也无法证明该用户有权执行哪些操作。
关于 AI 爬虫的常见问题
哪些 AI 爬虫会忽略 robots.txt?
根据运营方自己的文档,ChatGPT-User、Perplexity-User、Google-Agent、Google-GeminiNotebook、Amzn-User、meta-externalfetcher 和 facebookexternalhit 可能不会遵守 robots.txt,主要是因为抓取由用户触发。社区报告称 Bytespider 也不遵守该规则。
GPTBot、OAI-SearchBot 和 ChatGPT-User 有什么区别?
GPTBot 会收集网页,用于训练 OpenAI 模型;OAI-SearchBot 会为 ChatGPT 搜索建立网页索引;ChatGPT-User 则会在用户提出请求时抓取网页。前两者遵守 robots.txt;OpenAI 表示,其规则可能不适用于第三者。
哪些 AI 智能体没有 robots.txt token?
ChatGPT agent、Gemini Spark 和 Meta 的 Muse。可以通过来自 https://chatgpt.com 的 Web Bot Auth 签名验证 ChatGPT agent。Gemini Spark 在本地 Chrome 模式下使用的标识与用户自己的浏览器相同;Muse 则没有公开的 User-Agent。
如何验证请求确实来自 Googlebot?
Google 会以 JSON 文件公布 Googlebot 的 IP 地址段,链接可在此目录对应的条目中找到。请将请求的 IP 地址与这些地址段进行比对,不要只依赖 User-Agent 字符串。
这份 AI 爬虫列表来自哪里?
每个条目都链接到运营方的相关文档(如果有);缺少相关文档的 token 则来自社区维护的 ai.robots.txt 注册表。由于找不到运营方文档,我们未收录 DeepSeekBot、Kimi-User、Manus-User、NovaAct、Devin 和 PetalBot。此列表于 2026年10月1日核查。