Common Crawl · Training
What is CCBot?
Common Crawl's crawler, which builds an open web corpus that is widely used to train AI models. It follows robots.txt.
- robots.txt token
CCBot- Operator
- Common Crawl
- Purpose
- Training
- Obeys robots.txt
- Yes The operator says it follows robots.txt.
- Documented by operator
- Yes, on the operator's own pages.
- Published IP list
- index.commoncrawl.org/ccbot.json
- Documentation
- commoncrawl.org
Its hosts resolve to *.crawl.commoncrawl.org in reverse DNS, and Common Crawl publishes an IP list.
Block CCBot from your whole site
User-agent: CCBot
Disallow: /
Allow it, but keep private paths out
User-agent: CCBot
Allow: /
Disallow: /admin/
Add it to a robots.txt in the generator →