AI agent · August 2026
CCBot in robots.txt: who blocks it
7.4% of the 15,009 sites in the census name CCBot in their robots.txt, and 95.5% of those block it.
Sites naming it
- Sample
- n=15,009
- Coverage
- 100.0%
The August 2026 crawl, parsed for a robots.txt rule naming this user-agent, on the same 0.1% sample of the corpus as every other figure on this site — 15,009 sites. It used to be measured on the whole corpus and this note used to say so; it moved to the monthly sample when the axis gained a series, so it now carries the same margin as everything else.
Sites blocking it
- Sample
- n=15,009
- Coverage
- 100.0%
What CCBot actually does
Run by Common Crawl. Builds the public Common Crawl corpus, which many model trainers use. Blocking it is an indirect way of opting out of several models at once. Blocking it costs you no traffic: this crawler does not send anyone your way.
What to write in your robots.txt
Put the rule in its own group. A Disallow under User-agent: * predates these agents and several of them do not treat it as applying to them — which is also why this census does not count it.
User-agent: CCBot
Disallow: /User-agent: CCBot
Allow: /