---
title: "面向 AI Agent 的网站：llms.txt、robots.txt、AI 爬虫 · DotsAgent"
description: "生成 llms.txt、供 AI 爬虫使用的 robots.txt 和 RFC 9727 API 目录，还可按 token、用途和已公布的 IP 列表查询 40 种 AI 爬虫和 Agent。"
url: https://dotsagent.io/zh/agent-ready
---

开放你的网站 · 参考资料

# 让 AI Agent 能读懂你的网站

AI 爬虫和 Agent 读取网站的方式与人不同。下面的工具可以生成它们会查找的文件；后面的参考资料则介绍哪些约定已成为标准、哪些仍是草案，以及有多少客户端实际在使用。

- [llms.txt 生成器](https://dotsagent.io/zh/agent-ready/llms-txt): 输入名称、简介和链接分区，即可生成 llmstxt.org 格式的文件并复制。

- [AI 爬虫专用 robots.txt](https://dotsagent.io/zh/agent-ready/robots-txt): 按爬虫用途设置规则，单独启用或停用爬虫，添加 Content-Signal 并下载生成的文件。

- [API 目录](https://dotsagent.io/zh/agent-ready/api-catalog): 列出 API 及其 OpenAPI 文件和文档，即可生成 RFC 9727 linkset，以及用于提供该文件的 nginx 配置块。

- [AI 爬虫目录](https://dotsagent.io/zh/agent-ready/bots): 收录 40 种爬虫和 Agent，列出其 token、用途、robots.txt 行为和已公布的 IP 列表。

## 各项约定及其成熟度

截至 October 2026 的状态。每项都链接到规范或相关证据，方便你判断是否值得在自己的网站上采用。

### robots.txt 标准

网站根目录下的这个文件通过 token 指定爬虫，并列出各爬虫应跳过的路径；其规范为 RFC 9309。现在，AI 服务运营方会为训练、搜索和用户触发的抓取分别发布 token，因此可以通过同一个文件拒绝训练爬取，同时允许搜索爬取。遵守规则是自愿的，而且有些用户触发的抓取工具表示这些规则可能不适用于它们。

`/robots.txt`

```
User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /
```

[platform.openai.com](https://platform.openai.com/docs/bots)[developers.google.com](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers)

### Content-Signal 广泛使用

在 robots.txt 中额外添加一行，说明抓取的内容可用于哪些用途：search、ai-input 和 ai-train，每项设为 yes 或 no。Cloudflare 于 2025 年 9 月以 CC0 许可发布了这项政策，并通过其托管版 robots.txt 在约 380 万个域名上提供该政策。对应的 IETF 草案于 2026 年 4 月 4 日到期，而 IETF AIPREF 工作组仍在继续制定词汇表草案。

`/robots.txt`

```
User-agent: *
Content-Signal: ai-train=no, search=yes, ai-input=yes
Allow: /
```

[blog.cloudflare.com](https://blog.cloudflare.com/content-signals-policy/)[datatracker.ietf.org](https://datatracker.ietf.org/wg/aipref/documents/)

### llms.txt 提案

在 /llms.txt 提供一个 Markdown 文件，其中包含 H1 标题、引用摘要和带链接的 H2 章节，让 agent 无需解析导航栏即可访问关键页面。llmstxt.org 上的规范于 2026 年 8 月 10 日修订，llms-full.txt 不属于该规范。没有任何运营方声明主要 crawler 会读取此文件。

`/llms.txt`

```
# Acme API

> Card payments for small online shops.

## Docs

- [Quickstart](https://example.com/docs/quickstart.md): first request
```

[llmstxt.org](https://llmstxt.org/)[llmstxt.org](https://llmstxt.org/changes.html)

### 页面的 Markdown 版本 部分采用

为每个页面提供 Markdown 副本，路径可为 page.md 或 page.html.md；也可以在请求的 Accept 标头要求 text/markdown 时返回 Markdown。Checkly 在 2026 年 2 月的测试中发现，Claude Code、Cursor 和 OpenCode 会发送该标头，而 Codex、Gemini CLI、GitHub Copilot 和 Windsurf 不会。AnswerShare 统计了来自前沿 crawler 的 1,277,965 次请求，其中没有一次请求 Markdown。

`Request`

```
GET /docs/quickstart HTTP/1.1
Host: example.com
Accept: text/markdown
```

[developers.cloudflare.com](https://developers.cloudflare.com/fundamentals/reference/markdown-for-agents/)[checklyhq.com](https://www.checklyhq.com/blog/state-of-ai-agent-content-negotation/)[answershare.com](https://answershare.com/markdowndoesnotwork)

### API 目录 标准

RFC 9727 于 2025 年 6 月发布，规定使用 /.well-known/api-catalog 列出公开 API，并为每个 API 链接到机器可读的描述、文档和状态页。响应必须使用 application/linkset+json，其他页面则可以通过 api-catalog 链接关系指向该目录。

`/.well-known/api-catalog`

```
{
  "linkset": [
    {
      "anchor": "https://api.example.com/v1",
      "service-desc": [{ "href": "https://api.example.com/v1/openapi.json" }]
    }
  ]
}
```

[rfc-editor.org](https://www.rfc-editor.org/rfc/rfc9727.html)

### 签名的 agent 请求（Web Bot Auth） 草案

Agent 使用 HTTP Message Signatures（RFC 9421）为每个请求签名，并在 Signature-Agent 标头中标明其运营方。你可以根据运营方在 /.well-known/http-message-signatures-directory 下发布的密钥验证签名。IETF 工作组草案的日期为 2026 年 9 月 1 日；ChatGPT agent 已使用 https://chatgpt.com 进行签名，Google 正在测试 https://agent.bot.goog。

`Request header`

```
Signature-Agent: "https://chatgpt.com"
```

[datatracker.ietf.org](https://datatracker.ietf.org/doc/draft-ietf-webbotauth-httpsig-protocol/)[blog.cloudflare.com](https://blog.cloudflare.com/signed-agents/)

### NLWeb 停滞

Microsoft 的 NLWeb 为网站提供 /ask 和 /mcp endpoint，Agent 可通过自然语言查询；其规范版本为 0.55。Lumar 报告称，nlweb.ai 的证书已于 4 August 2026 过期，且该 repository 自六月以来一直没有动静，因此在基于它构建前请先观望。

[github.com](https://github.com/microsoft/NLWeb/blob/main/docs/nlweb-rest-api.md)[agentic-readiness.lumar.io](https://agentic-readiness.lumar.io/docs/capabilities/nl-web)

## 不主动表明身份的 Agent

有些 Agent 不会发布 robots.txt token，而是像用户一样浏览，因此任何 User-agent 规则都无法匹配它们。ChatGPT agent 使用 Web Bot Auth 为请求签名，并发送 Signature-Agent: "https://chatgpt.com"；你可以通过 OpenAI 的密钥目录验证该签名。Cloudflare 将这类流量标记为 chatgpt-agent。Gemini Spark 的本地 Chrome 模式和 Meta 的 Muse 均没有任何公开文档说明的标识。

## 关于适配 Agent 的网站

### 什么样的网站适配 Agent？

目前没有单一标准对此作出定义。通常，这意味着 robots.txt 按用途列出 AI bot，可选提供 llms.txt 和 API 目录等文件，为关键页面提供 Markdown 版本，并提供验证获准访问的 Agent 的方式。在这些做法中，只有 robots.txt 和 API 目录是已发布的 RFC。

### AI crawler 会读取 llms.txt 吗？

我们没有找到任何运营方文档说明目录中的 crawler 会获取该文件。Lighthouse 会检查文件是否存在，而你指定的 Agent 可以像读取其他页面一样读取它。应将其视为方便 Agent 使用的文件，而不是搜索功能。

### 我应该在网站上屏蔽 AI crawler 吗？

请按用途分别决定，而不是一概而论。你可以拒绝 GPTBot、ClaudeBot 和 Google-Extended 等训练 token，而不影响搜索。屏蔽 OAI-SearchBot 或 PerplexityBot 等搜索 token，则会让你的网站完全无法出现在相应产品中。

### 可以通过 robots.txt 屏蔽 ChatGPT agent 吗？

不可以，因为它没有发布 token。你可以根据来自 https://chatgpt.com 的 Web Bot Auth 签名识别它，并在服务器或 CDN 上允许或拒绝请求。有效签名只能证明请求由哪个运营方发出，不能证明用户身份或用户有权执行哪些操作。

### Content-Signal 是官方标准吗？

不是。它是 Cloudflare 于 September 2025 推出的 CC0 策略，通过 Cloudflare 管理的 robots.txt 部署在约 3.8 million 个域名上。它的 IETF 草案已于 4 April 2026 过期，AIPREF 工作组仍在起草通用词汇表。

为 AI agent 开发者提供的独立参考资料。与此处提及的任何厂商均无关联。

© 2026 DotsAgent · 事实核查日期：2026年10月1日
