7.0%
Sites naming it

Sites naming it

Sample
n=15,009
Coverage
100.0%

The August 2026 crawl, parsed for a robots.txt rule naming this user-agent, on the same 0.1% sample of the corpus as every other figure on this site — 15,009 sites. It used to be measured on the whole corpus and this note used to say so; it moved to the monthly sample when the axis gained a series, so it now carries the same margin as everything else.

6.5%
Sites blocking it

Sites blocking it

Sample
n=15,009
Coverage
100.0%
Name it
Name it: 7.0%
7.0%
Block it
Block it: 6.5%
6.5%
00.250.500.751.0
Of the 15,009 sites in the census, 7.0% name Google-Extended and 6.5% forbid it. Both bars are the same base, so the gap between them is the sites that named it to let it in.

What Google-Extended actually does

Run by Google. Controls whether your content is used to train Google’s generative models. It does NOT affect Google Search: blocking it leaves your search ranking alone. Blocking it costs you no traffic: this crawler does not send anyone your way.

What to write in your robots.txt

Put the rule in its own group. A Disallow under User-agent: * predates these agents and several of them do not treat it as applying to them — which is also why this census does not count it.

User-agent: Google-Extended
Disallow: /
robots.txtTurns Google-Extended away. Blocking it costs you no traffic: this crawler does not send anyone your way.
User-agent: Google-Extended
Allow: /
robots.txtLets it through, said out loud rather than by saying nothing. Either way it is a published preference and not a wall: robots.txt enforces nothing.