AIVIS

AI Crawler Blocking Report 2026

Which AI crawlers do websites shut out, and which do they let in? Measured on 1,349 real robots.txt files, not a survey.

Sample: 1,349 robots.txt files, out of 1,769 domains scanned · Scans from Sep 19, 2026 to Sep 22, 2026 · Recomputed Sep 25, 2026

11.4%Block at least one AI crawler
7.6%Block GPTBot
1.4%Block all 13 AI crawlers
1,349robots.txt files analysed

11.4% of the 1,349 robots.txt files AIVIS analysed shut at least one AI crawler out of the whole site, and 1.4% shut out all 13 crawlers tracked. The large majority leave every AI crawler free to read them.

Blocking targets model training far more than answers: 11.4% of sites block at least one training crawler, while only 5.8% block an AI search crawler and 5.6% block the fetchers an assistant sends when a user asks it to open a page. 5.1% block training crawlers only, keeping AI search and live fetches open.

Among sectors with enough domains to be representative, News & Media blocks most often (50.9% of robots.txt files block at least one AI crawler) and Universities least often (0%).

Blocking rate by crawler

Share of the 1,349 robots.txt files that disallow each crawler from the whole site.

Crawler (user-agent)OperatorPurposeBlocked by
CCBotCommon CrawlModel training8.9%
BytespiderByteDanceModel training8.6%
GPTBotOpenAIModel training7.6%
ClaudeBotAnthropicModel training7.3%
Google-ExtendedGoogleModel training6.7%
anthropic-aiAnthropicModel training6.6%
Applebot-ExtendedAppleModel training6.6%
Meta-ExternalAgentMetaModel training6.5%
PerplexityBotPerplexityAI search index5.6%
ChatGPT-UserOpenAIUser-requested fetch4.1%
Perplexity-UserPerplexityUser-requested fetch3.9%
Claude-UserAnthropicUser-requested fetch3.6%
OAI-SearchBotOpenAIAI search index3%

AI crawler blocking by sector

Share of robots.txt files in each industry that block at least one AI crawler, and the four crawlers asked about most.

Sectorrobots.txt filesAny AI crawlerGPTBotClaudeBotPerplexityBotGoogle-Extended
All websites1,34911.4%7.6%7.3%5.6%6.7%
SaaS & software1197.6%0.8%0.8%0%0.8%
News & Media5750.9%38.6%35.1%31.6%33.3%
Universities460%0%0%0%0%
E-commerce3611.1%8.3%5.6%0%8.3%
Banking & Finance2516%8%4%4%8%
Public Sector238.7%8.7%4.3%8.7%4.3%
Hotels & B&Bs128.3%8.3%8.3%0%8.3%
Automotive1020%20%20%10%20%
Telecom & Utilities(low sample)80%0%0%0%0%
Restaurants(low sample)850%25%37.5%25%37.5%
Travel & Transport(low sample)633.3%0%0%0%0%
Real estate agencies(low sample)616.7%16.7%0%0%0%
Web agencies(low sample)40%0%0%0%0%
Law firms & professional services(low sample)20%0%0%0%0%
Local businesses & tradespeople(low sample)20%0%0%0%0%

Sectors are assigned by the AIVIS classifier from homepage content: 390 of the 1,769 domains have one, the rest count in the totals but not in this table. Sectors with fewer than 10 robots.txt files are shown greyed out and marked low_sample in the dataset: too few to represent the industry.

Methodology

Download the dataset

Totals, the rate for every crawler and every sector row behind this report, recomputed live as the corpus grows.

Licensed under CC-BY 4.0 — free to reuse with attribution to AIVIS.

Cite as: AIVIS, AI Crawler Blocking Report 2026, https://aivis.lumnika.com/en/research/ai-crawler-blocking-2026

Is your site blocking the AI crawlers you need?

Check your robots.txt against all 13 AI crawlers in seconds, for free, and see exactly which rule shuts each one out.