---
title: "CCBot: block or allow this AI crawler · DotsAgent"
description: "CCBot is run by Common Crawl and listed under training. Its robots.txt token, whether it follows the rules, its IP list and how to control it."
url: https://dotsagent.io/agent-ready/bots/ccbot
---

Common Crawl · Training

# What is CCBot?

Common Crawl's crawler, which builds an open web corpus that is widely used to train AI models. It follows robots.txt.

- **robots.txt token**: `CCBot`
- **Operator**: Common Crawl
- **Purpose**: Training
- **Obeys robots.txt**: The operator says it follows robots.txt.
- **Documented by operator**: Yes, on the operator's own pages.
- **Published IP list**: [index.commoncrawl.org/ccbot.json](https://index.commoncrawl.org/ccbot.json)
- **Documentation**: [commoncrawl.org](https://commoncrawl.org/ccbot)

Its hosts resolve to *.crawl.commoncrawl.org in reverse DNS, and Common Crawl publishes an IP list.

## Block CCBot from your whole site

`robots.txt`

```
User-agent: CCBot
Disallow: /
```

## Allow it, but keep private paths out

`robots.txt`

```
User-agent: CCBot
Allow: /
Disallow: /admin/
```

[Add it to a robots.txt in the generator →](https://dotsagent.io/agent-ready/robots-txt)

## Sources

1. [commoncrawl.org](https://commoncrawl.org/ccbot)/ccbot

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026
