AI Crawler Blocking Report 2026
Which AI crawlers do websites shut out, and which do they let in? Measured on 1,349 real robots.txt files, not a survey.
Sample: 1,349 robots.txt files, out of 1,769 domains scanned · Scans from Sep 19, 2026 to Sep 22, 2026 · Recomputed Sep 25, 2026
11.4% of the 1,349 robots.txt files AIVIS analysed shut at least one AI crawler out of the whole site, and 1.4% shut out all 13 crawlers tracked. The large majority leave every AI crawler free to read them.
Blocking targets model training far more than answers: 11.4% of sites block at least one training crawler, while only 5.8% block an AI search crawler and 5.6% block the fetchers an assistant sends when a user asks it to open a page. 5.1% block training crawlers only, keeping AI search and live fetches open.
Among sectors with enough domains to be representative, News & Media blocks most often (50.9% of robots.txt files block at least one AI crawler) and Universities least often (0%).
Blocking rate by crawler
Share of the 1,349 robots.txt files that disallow each crawler from the whole site.
| Crawler (user-agent) | Operator | Purpose | Blocked by |
|---|---|---|---|
CCBot | Common Crawl | Model training | 8.9% |
Bytespider | ByteDance | Model training | 8.6% |
GPTBot | OpenAI | Model training | 7.6% |
ClaudeBot | Anthropic | Model training | 7.3% |
Google-Extended | Model training | 6.7% | |
anthropic-ai | Anthropic | Model training | 6.6% |
Applebot-Extended | Apple | Model training | 6.6% |
Meta-ExternalAgent | Meta | Model training | 6.5% |
PerplexityBot | Perplexity | AI search index | 5.6% |
ChatGPT-User | OpenAI | User-requested fetch | 4.1% |
Perplexity-User | Perplexity | User-requested fetch | 3.9% |
Claude-User | Anthropic | User-requested fetch | 3.6% |
OAI-SearchBot | OpenAI | AI search index | 3% |
AI crawler blocking by sector
Share of robots.txt files in each industry that block at least one AI crawler, and the four crawlers asked about most.
| Sector | robots.txt files | Any AI crawler | GPTBot | ClaudeBot | PerplexityBot | Google-Extended |
|---|---|---|---|---|---|---|
| All websites | 1,349 | 11.4% | 7.6% | 7.3% | 5.6% | 6.7% |
| SaaS & software | 119 | 7.6% | 0.8% | 0.8% | 0% | 0.8% |
| News & Media | 57 | 50.9% | 38.6% | 35.1% | 31.6% | 33.3% |
| Universities | 46 | 0% | 0% | 0% | 0% | 0% |
| E-commerce | 36 | 11.1% | 8.3% | 5.6% | 0% | 8.3% |
| Banking & Finance | 25 | 16% | 8% | 4% | 4% | 8% |
| Public Sector | 23 | 8.7% | 8.7% | 4.3% | 8.7% | 4.3% |
| Hotels & B&Bs | 12 | 8.3% | 8.3% | 8.3% | 0% | 8.3% |
| Automotive | 10 | 20% | 20% | 20% | 10% | 20% |
| Telecom & Utilities(low sample) | 8 | 0% | 0% | 0% | 0% | 0% |
| Restaurants(low sample) | 8 | 50% | 25% | 37.5% | 25% | 37.5% |
| Travel & Transport(low sample) | 6 | 33.3% | 0% | 0% | 0% | 0% |
| Real estate agencies(low sample) | 6 | 16.7% | 16.7% | 0% | 0% | 0% |
| Web agencies(low sample) | 4 | 0% | 0% | 0% | 0% | 0% |
| Law firms & professional services(low sample) | 2 | 0% | 0% | 0% | 0% | 0% |
| Local businesses & tradespeople(low sample) | 2 | 0% | 0% | 0% | 0% | 0% |
Sectors are assigned by the AIVIS classifier from homepage content: 390 of the 1,769 domains have one, the rest count in the totals but not in this table. Sectors with fewer than 10 robots.txt files are shown greyed out and marked low_sample in the dataset: too few to represent the industry.
Methodology
- Sample: the latest scan of every domain in the AIVIS corpus (1,769 domains, scanned between Sep 19, 2026 and Sep 22, 2026). The base is the 1,349 domains that serve a robots.txt (76.3% of the corpus): a site without one sets no restriction at all, and is left out of the percentages rather than counted as open.
- Measure: for each of the 13 user-agent tokens, the robots.txt is evaluated with standard group matching (RFC 9309) — the group naming the token, or the catch-all “User-agent: *” group when none does. A crawler counts as blocked when the site root “/” is disallowed for it. Rules that close only some folders do not count as a block.
- Purpose follows each operator’s own documentation: training crawlers collect data for models (GPTBot, ClaudeBot, anthropic-ai, Google-Extended, Applebot-Extended, CCBot, Bytespider, Meta-ExternalAgent); AI search crawlers build the index answers cite (OAI-SearchBot, PerplexityBot); user-requested fetchers open a page when a person asks the assistant to (ChatGPT-User, Claude-User, Perplexity-User). anthropic-ai is a legacy token that many robots.txt files still list.
- Limits: robots.txt is a declaration, not enforcement. Firewalls and CDN bot protection can refuse crawlers that robots.txt allows, and some crawlers ignore it; this report measures the declared policy only.
- Relation to the AI Readiness Index: the Index reports blocking over all scanned domains, including those without a robots.txt; this report uses robots.txt files as its base, so its percentages are higher by definition. Both are recomputed live from the same scans.
Download the dataset
Totals, the rate for every crawler and every sector row behind this report, recomputed live as the corpus grows.
Licensed under CC-BY 4.0 — free to reuse with attribution to AIVIS.
Cite as: AIVIS, AI Crawler Blocking Report 2026, https://aivis.lumnika.com/en/research/ai-crawler-blocking-2026
Is your site blocking the AI crawlers you need?
Check your robots.txt against all 13 AI crawlers in seconds, for free, and see exactly which rule shuts each one out.