AI Crawler Access: Which Bots to Allow or Block in 2026
A practical walkthrough of what the AI Crawler Access Checker shows you, why "block all AI bots" and "allow all AI bots" are both too blunt for most sites, and the vendor-documented split between training crawlers, search crawlers and user-triggered fetchers.
Why "Block All AI Bots" Is the Wrong Question
Most site owners who go looking for an AI-bot robots.txt snippet find a single wildcard block: disallow GPTBot, disallow CCBot, disallow Bytespider, done. That answers one question — "can AI companies use my content to train future models?" — while silently answering a completely different question the same way: "can my content show up when someone asks ChatGPT, Claude or Perplexity a question?" Those are not the same decision, and in 2026 the major AI vendors run separate bots specifically so you don't have to make them together.
OpenAI's own crawler documentation is explicit about this: OAI-SearchBot indexes pages for ChatGPT Search, ChatGPT-User fetches a page only when a person asks ChatGPT to open it, and GPTBot is the separate crawler that collects content for possible model training. You can allow the first two and block the third. Anthropic runs the same three-way split for Claude — Claude-SearchBot, Claude-User and ClaudeBot — and Anthropic's public crawler documentation confirms they're independent agents that each respect their own robots.txt rule.
The Three Jobs an AI Crawler Can Have
Once you group crawlers by job instead of by company, the picture gets much simpler. Every major AI vendor's bot fleet falls into one of three categories:
| Job | What it does | Examples | Blocking it means |
|---|---|---|---|
| Search / answer crawler | Indexes or retrieves pages specifically to cite in live AI answers | OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Applebot | You stop appearing in that product's AI-generated answers |
| User-triggered fetcher | Fetches one page the moment a person asks the assistant to open it | ChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher | Often unreliable anyway — several vendors say these may not fully honor robots.txt since a human explicitly requested the page |
| Model-training crawler | Collects content that may be used to train a future model | GPTBot, ClaudeBot, meta-externalagent, Google-Extended, Applebot-Extended | Usually nothing to your current search or AI-answer visibility — it's a separate, forward-looking policy call |
The AI Crawler Access Checker's three sections — search & assistant crawlers, model-training crawlers, and other data-scraping bots — follow this same split, and the "Recommended" preset in the generator reflects the position most publishers land on: stay open to search and answer crawlers, and treat training crawlers as an optional, separate decision.
The Google-Extended Exception
Google's setup deserves its own callout because it doesn't follow the same clean split. There is no separate "Google AI search" bot — regular Googlebot powers both classic Search results and AI Overviews, so blocking Googlebot removes you from both at once. Google-Extended is a different kind of token entirely: it doesn't crawl your site itself, it controls whether content that Googlebot already crawled can be used to train future Gemini models and support some AI grounding features. Google states plainly that Google-Extended has no effect on Search ranking or AI Overviews eligibility, so blocking it is a clean way to opt out of the training use case without touching your Search visibility either way.
Apple's setup is closer to OpenAI and Anthropic's: Applebot handles Siri, Spotlight and Apple's AI answers, while Applebot-Extended is the separate control token for whether that same crawled data may train Apple's models. Block the -Extended token, keep the base bot allowed, and you get the same outcome — visible in the product, opted out of training.
How to Read the Checker's Results
Paste your live robots.txt content (not a URL — the actual file text) into the checker and it evaluates all 19 bots it tracks against your rules the way a crawler actually would: it looks for a rule group that names that exact bot first, and only falls back to your general User-agent: * group if no specific rule exists for it. That matters because a lot of sites assume a blanket "block everything under *" rule also blocks named AI bots — it does, but only until you add an explicit Allow line for the ones you actually want in.
Each bot gets one of four statuses. Allowed means the matching rule group has no root-level disallow (or has an explicit Allow: /). Blocked means a Disallow: / applies with no overriding Allow: /. Partial means only specific paths are disallowed, not the whole site. Not mentioned means the bot has no dedicated rule and your file has no User-agent: * group either, so it's open to the web by default — the same as any browser.
The checker deliberately keeps this simple rather than modeling every edge case in the Robots Exclusion Protocol (RFC 9309) — things like longest-path-wins precedence for deeply nested Allow/Disallow rules on the same bot. For the common case most sites actually have (a handful of Disallow paths, maybe one bot-specific block), it gives you an accurate, honest picture. For anything unusual, read the raw rule group yourself.
An Honest Limit: Robots.txt Is a Request
Nothing about a Disallow line stops a bot from ignoring it. Reputable AI vendors say they respect robots.txt for their crawling bots, but they're explicit that the picture is murkier for user-triggered fetchers: OpenAI's documentation notes ChatGPT-User requests may not always be subject to robots.txt checks, and Perplexity states outright that Perplexity-User generally does not check it, since the request only happens because a person is actively looking at your page through the assistant in that moment. Less-documented scrapers like Bytespider are widely reported to disregard robots.txt entirely in practice. If keeping a bot out is essential — not just a stated preference — you need to enforce it at the CDN, firewall or WAF layer using the vendor's published IP ranges, not a text file that operates on the honor system.
It's also worth knowing that major AI vendors add new crawler tokens and split existing ones into narrower roles several times a year, so a policy you set once will drift out of date. Re-run the checker every few months and cross-reference against the vendors' own current documentation — linked in the tool's help panel — since that's always the source of truth over any third-party list, including this one.
Using the Generator
Once you've decided on a policy, the second half of the tool builds an explicit robots.txt block: one User-agent / Allow or Disallow pair per bot, grouped and commented by category. Start from the "Recommended" preset, flip any individual bot you disagree with, and paste the result below your site's existing general rules — it's designed to be appended, not to replace your whole robots.txt. Re-run the checker afterward to confirm the new block reads the way you intended.
Related tools: Robots & Sitemap Generator · llms.txt Generator · AI Visibility Checker · Technical SEO & Sitemaps Hub