← Back to all tools

Technical SEO & Sitemaps

AI Crawler Access Checker

Paste your robots.txt and see exactly which AI crawlers — ChatGPT, Claude, Perplexity, Google, Meta, Apple and more — can reach your site, which are blocked, and which you've never addressed at all. Then generate a ready-to-paste AI-bot policy block.

1. Check your current robots.txt

Paste your live robots.txt content below (not the URL — copy the actual file contents from yoursite.com/robots.txt). Nothing is uploaded; this runs entirely in your browser.

Checks against 19 documented AI crawler user-agents.
⚠️ Robots.txt is a request, not a lock. OpenAI says ChatGPT-User may not always honor it, and Perplexity says Perplexity-User generally ignores it entirely, since both fire only after a person explicitly asks the assistant to open your page. Treat a "blocked" status here as your stated policy, not a guarantee — enforce it at your CDN or firewall if it's essential. Google-Extended and Applebot-Extended are policy tokens, not bots that visit your logs — don't expect to see them as requests.

🔍 Search & assistant crawlers

Bots like OAI-SearchBot, Claude-SearchBot and PerplexityBot index or fetch pages specifically to answer live questions and cite sources. Blocking these directly removes you from AI-generated answers — the opposite of what most sites want in 2026.

🎓 Model-training crawlers

Bots like GPTBot, ClaudeBot and meta-externalagent collect content that may train future models. This is a separate policy decision from search visibility — you can block training while staying fully visible in AI answers, since most vendors run independent bots for each job.

🧩 They're not interchangeable

OpenAI explicitly lets you allow OAI-SearchBot while disallowing GPTBot — one rule doesn't control both. The same split exists at Anthropic (Claude-SearchBot/Claude-User vs. ClaudeBot). A single wildcard "block all AI" rule can't express this nuance.

🇬 The Google-Extended exception

Google states Google-Extended does not affect Search ranking or AI Overviews eligibility — it only controls whether crawled content can help train future Gemini models and some grounding uses. Blocking it doesn't hurt your Google visibility either way.

🗂️ Other scraper bots

CCBot (Common Crawl), Bytespider (ByteDance), Amazonbot, Diffbot and cohere-ai are less formally documented than the big six vendors' bots, and several are reported to crawl aggressively or ignore robots.txt outright — treat a robots.txt rule against them as a weak signal, not a guarantee.

🔄 Policies change often

New AI crawler tokens appear several times a year as vendors launch products or split existing bots into narrower roles. Re-run this checker every few months and compare against the vendors' own crawler documentation, which is always the source of truth.

2. Generate an AI crawler policy block

Pick a starting policy, adjust individual bots if you want, then generate an explicit robots.txt block. Paste it into your existing robots.txt below your general User-agent: * rules — don't replace your whole file with just this.

Checked = allowed. Unchecked = disallowed.