The question isn't "block all AI bots or allow all of them": different bots do different things, and the right choice depends on your goal.
Two categories, two different decisions
Training bots (GPTBot, ClaudeBot, CCBot, Bytespider, Meta-ExternalAgent): gather content to train future models. Blocking them doesn't exclude you from today's answers, but from the data future models get trained on.
Citation / live browsing bots (OAI-SearchBot, Perplexity-User, ChatGPT-User): used in real time to answer a specific query with a citation. Blocking these excludes the site from AI answers **starting today**.
The most common choice (and our default recommendation)
If your goal is to be cited and get referral traffic from AI assistants (which converts roughly 5x better than organic, 2026 market data), allowing both categories is almost always the right call: the more training sources know about you, the more likely you are to be cited in the future.
When blocking makes sense
- Exclusively licensed or paid content you don't want turned into training material without compensation.
- Restricted areas (user dashboards, paywalled content) — but these should already be excluded for any crawler, not just AI ones.
Ready-made template
We've published an annotated AI-friendly robots.txt template that explicitly allows all known AI bots while keeping your existing rules. Or, an AIVIS report generates the corrected version for your specific domain.