dotsagent.io
语言:简体中文
40 个爬虫和代理

AI 爬虫和代理目录

这里列出了我们能够确认的所有 AI 机器人,包括 robots.txt 中使用的令牌、抓取目的、运营方是否声明遵守相关规则,以及如何验证请求确实来自该机器人。令牌仅见于社区维护的 ai.robots.txt 登记表时,会在对应页面特别标注。

机器人令牌用途遵守 robots.txt验证
GPTBot
OpenAI
GPTBot训练是IP 列表
OAI-SearchBot
OpenAI
OAI-SearchBot搜索是IP 列表
ChatGPT-User
OpenAI
ChatGPT-User用户请求否IP 列表
OAI-AdsBot
OpenAI
OAI-AdsBot广告未说明IP 列表
ChatGPT agent
OpenAI
无令牌Agent未说明签名请求
ClaudeBot
Anthropic
ClaudeBot训练是IP 列表
Claude-SearchBot
Anthropic
Claude-SearchBot搜索是IP 列表
Claude-User
Anthropic
Claude-User用户请求是IP 列表
Googlebot
Google
Googlebot搜索是IP 列表
Google-Extended
Google
Google-Extended训练是—
Google-CloudVertexBot
Google
Google-CloudVertexBot搜索是IP 列表
GoogleOther
Google
GoogleOther训练是IP 列表
Google-Agent
Google
Google-AgentAgent否IP 列表
Google-GeminiNotebook
Google
Google-GeminiNotebook用户请求否IP 列表
Gemini Spark
Google
无令牌Agent未说明—
Gemini-Deep-Research
Google
Gemini-Deep-Research用户请求未说明—
GoogleAgent-Mariner
Google
GoogleAgent-MarinerAgent未说明—
Applebot
Apple
Applebot搜索是IP 列表
Applebot-Extended
Apple
Applebot-Extended训练是—
PerplexityBot
Perplexity
PerplexityBot搜索是IP 列表
Perplexity-User
Perplexity
Perplexity-User用户请求否IP 列表
meta-externalagent
Meta
meta-externalagent训练是—
meta-webindexer
Meta
meta-webindexer搜索是—
meta-externalfetcher
Meta
meta-externalfetcher用户请求否—
meta-externalads
Meta
meta-externalads广告是—
facebookexternalhit
Meta
facebookexternalhit用户请求否—
Muse
Meta
无令牌Agent未说明—
Amazonbot
Amazon
Amazonbot训练是IP 列表
Amzn-SearchBot
Amazon
Amzn-SearchBot搜索是IP 列表
Amzn-User
Amazon
Amzn-User用户请求否IP 列表
DuckAssistBot
DuckDuckGo
DuckAssistBot搜索是IP 列表
MistralAI-User
Mistral
MistralAI-User用户请求是IP 列表
MistralAI-Index
Mistral
MistralAI-Index搜索是IP 列表
MistralAI-Training
Mistral
MistralAI-Training训练是—
CCBot
Common Crawl
CCBot训练是IP 列表
cohere-ai
Cohere
cohere-ai用户请求未说明—
Bytespider
ByteDance
Bytespider训练否—
YouBot
You.com
YouBot搜索未说明—
Diffbot
Diffbot
Diffbot搜索未说明—
Timpibot
Timpi
Timpibot训练未说明—

用途说明

训练
收集网页,用于模型训练或一般研究爬取。Google-Extended 和 Applebot-Extended 是控制令牌,本身不会抓取任何内容。
搜索
为搜索产品编入网页索引,涵盖 Google Search 以及 ChatGPT search、Perplexity 和 DuckAssist 等 AI 问答引擎。屏蔽此类机器人,就无法出现在对应产品中。
用户请求
根据用户在聊天或工具中的请求抓取网页。多家运营方表示,robots.txt 可能不适用于这类请求。
Agent
代表用户浏览网页并执行操作,通常使用完整浏览器。许多代理没有令牌,因此只能通过签名识别(如果能够识别的话)。
广告
检查广告主提交的广告落地页。这些机器人不会包含在 robots.txt 生成器中。

验证爬虫身份

伪造 User-Agent 字符串毫不费力。大多数大型运营方会以 JSON 文件公布其爬虫使用的 IP 地址段,表格中标记为「IP 列表」。将请求的 IP 地址与该 token 对应的列表进行比对;仅凭名称匹配不能视为已验证。

部分运营方还会说明反向 DNS 信息。例如,Common Crawl 表示 CCBot 主机名会解析到 *.crawl.commoncrawl.org;在信任该主机前,请确认该主机名解析回同一个 IP 地址。

越来越多的智能体依据 Web Bot Auth 草案,使用 HTTP Message Signatures (RFC 9421) 对请求进行签名。请求通过 Signature-Agent 声明其运营方,你可以使用该域名下 /.well-known/http-message-signatures-directory 中的公钥验证签名。有效签名只能证明运营方身份,无法证明背后的用户是谁,也无法证明该用户有权执行哪些操作。

根据此列表生成 robots.txt →

关于 AI 爬虫的常见问题

哪些 AI 爬虫会忽略 robots.txt?

根据运营方自己的文档,ChatGPT-User、Perplexity-User、Google-Agent、Google-GeminiNotebook、Amzn-User、meta-externalfetcher 和 facebookexternalhit 可能不会遵守 robots.txt,主要是因为抓取由用户触发。社区报告称 Bytespider 也不遵守该规则。

GPTBot、OAI-SearchBot 和 ChatGPT-User 有什么区别?

GPTBot 会收集网页,用于训练 OpenAI 模型;OAI-SearchBot 会为 ChatGPT 搜索建立网页索引;ChatGPT-User 则会在用户提出请求时抓取网页。前两者遵守 robots.txt;OpenAI 表示,其规则可能不适用于第三者。

哪些 AI 智能体没有 robots.txt token?

ChatGPT agent、Gemini Spark 和 Meta 的 Muse。可以通过来自 https://chatgpt.com 的 Web Bot Auth 签名验证 ChatGPT agent。Gemini Spark 在本地 Chrome 模式下使用的标识与用户自己的浏览器相同;Muse 则没有公开的 User-Agent。

如何验证请求确实来自 Googlebot?

Google 会以 JSON 文件公布 Googlebot 的 IP 地址段,链接可在此目录对应的条目中找到。请将请求的 IP 地址与这些地址段进行比对,不要只依赖 User-Agent 字符串。

这份 AI 爬虫列表来自哪里?

每个条目都链接到运营方的相关文档(如果有);缺少相关文档的 token 则来自社区维护的 ai.robots.txt 注册表。由于找不到运营方文档,我们未收录 DeepSeekBot、Kimi-User、Manus-User、NovaAct、Devin 和 PetalBot。此列表于 2026年10月1日核查。

为 AI agent 开发者提供的独立参考资料。与此处提及的任何厂商均无关联。

© 2026 DotsAgent · 事实核查日期:2026年10月1日