AI crawler blocking and llms.txt statistics
How many websites block AI crawlers such as GPTBot, and how many publish llms.txt? This page collects the published figures in one place, with a source for each. None of these numbers are Aveena’s own measurements.
Key figures
News publishers block AI crawlers most
Large news sites were the first to block AI crawlers. By the end of 2023, 48% of the most widely used news sites across ten countries, including the UK, were blocking OpenAI’s crawlers, and 24% were blocking Google’s AI crawler.[1] Among 106 leading UK, US and world news sites, 56.6% blocked GPTBot, about a quarter blocked Google-Extended, and 42.5% blocked no AI crawlers at all.[2]
- GPTBotUK, US and world news (Press Gazette)56.6%
- OpenAI crawlers10 countries (Reuters Institute)48%
- Google AI crawler10 countries (Reuters Institute)24%
| Measure | Sample | Share |
|---|---|---|
| GPTBot | UK, US and world news (Press Gazette)[2] | 56.6% |
| OpenAI crawlers | 10 countries (Reuters Institute)[1] | 48% |
| Google AI crawler | 10 countries (Reuters Institute)[1] | 24% |
Most websites don't block AI crawlers
Outside news, blocking is far less common. Only about 37% of the top 10,000 domains on Cloudflare have a robots.txt file at all. Of those files, 7.8% disallow GPTBot and 5.6% disallow Google-Extended, while ClaudeBot, PerplexityBot and anthropic-ai each appear in under 5%.[3]
- GPTBotOpenAI7.8%
- Google-ExtendedGoogle5.6%
| Measure | Sample | Share |
|---|---|---|
| GPTBot | OpenAI[3] | 7.8% |
| Google-Extended | Google[3] | 5.6% |
About one site in ten publishes llms.txt
llms.txt adoption sits at around 8-10% however the sample is drawn: 10.13% of about 300,000 domains in late 2025,[4] and 9.3% of the top 1,000 and 8.3% of the top 10,000 domains in September 2026.[5] The same SE Ranking study found no correlation between having an llms.txt file and how often a domain is cited by AI.[4]
- Top 1,000 domainsRankability, Sep 20269.3%
- Top 10,000 domainsRankability, Sep 20268.3%
- About 300,000 domainsSE Ranking, Nov 202510.13%
| Measure | Sample | Share |
|---|---|---|
| Top 1,000 domains | Rankability, Sep 2026[5] | 9.3% |
| Top 10,000 domains | Rankability, Sep 2026[5] | 8.3% |
| About 300,000 domains | SE Ranking, Nov 2025[4] | 10.13% |
Why the figures differ
The studies measure different things. News-site studies look at a few dozen large publishers, which have strong reasons to block AI training. Cloudflare’s figure is a share of robots.txt files among top domains, and most of those domains have no robots.txt. Compare figures only within the same kind of sample.
What we don't know yet: UK businesses
There is no large published study of UK business websites. The UK appears in the news-site research above, but those are national publishers, not typical businesses. Until that data exists, the best way to know where your own site stands is to check it.
Check your own site
The free AI visibility checker reads your robots.txt for each AI crawler and checks for llms.txt. To change your rules, use the robots.txt generator, or build a file with the llms.txt generator.
Sources
- Richard Fletcher, How many news websites block AI crawlers?, Reuters Institute for the Study of Journalism, 22 February 2024. Sample: The 15 most-used online news sites in each of 10 countries, including the UK, at the end of 2023.
- Bron Maher, Which of the top 100 UK and US news websites are blocking AI crawlers, Press Gazette, 27 February 2024. Sample: 106 sites from Press Gazette's top 50 rankings for UK, US and world news.
- Reid Tatoris, Control content use for AI training with Cloudflare's managed robots.txt, Cloudflare, 1 July 2025. Sample: Robots.txt files of the top 10,000 domains on Cloudflare.
- Yulia Deda, LLMs.txt: why brands rely on it and why it doesn't work, SE Ranking, 7 November 2025. Sample: About 300,000 domains.
- llms.txt adoption, Rankability, 18 September 2026. Sample: Top 1,000 and top 10,000 domains on the Tranco list, collected 18 September 2026.
Figures checked against each source on 28 September 2026.