Your robots.txt is blocking ChatGPT without you knowing

Cyril Cretin
Founder RadarLLM.ai
Your website may be invisible to ChatGPT, Claude, Perplexity AND Gemini right now. The most common cause: your robots.txt. Many hosting providers, CDNs and security plugins block AI crawlers by default, without telling you. The result: AIs simply cannot read your site.
5.89% of all websites block GPTBot. And among major publishers, 79% block at least one AI bot. If you've never checked your robots.txt, there's a strong chance you're part of these statistics.
AI bots you need to know
Each AI platform uses one or more bots to crawl the web. Here is the complete list of bots your robots.txt must allow:
| Bot | Platform | Purpose | Allow? |
|---|---|---|---|
| GPTBot | ChatGPT | Training + Search | Yes |
| OAI-SearchBot | ChatGPT Search | Real-time search | Yes |
| ChatGPT-User | ChatGPT | User browsing | Yes |
| ClaudeBot | Claude | Training + citations | Yes |
| anthropic-ai | Anthropic | Training | Yes |
| PerplexityBot | Perplexity | Search + citations | Yes |
| Google-Extended | Gemini | Training + AI Overviews | Yes |
| Googlebot | Standard search | Yes |
How to check if your robots.txt blocks AI
Open your browser and go to your-site.com/robots.txt
Search for each bot name listed above (GPTBot, ClaudeBot, PerplexityBot, Google-Extended...)
Look for "Disallow: /" rules associated with these bots — that's the blocking signal
Here are the most common patterns that block AI crawlers:
# Examples of rules that BLOCK AI crawlers:
User-agent: GPTBot
Disallow: / # ← Blocks ChatGPT
User-agent: ClaudeBot
Disallow: / # ← Blocks Claude
User-agent: *
Disallow: / # ← Blocks EVERYTHING, including AI
User-agent: CCBot
Disallow: / # ← Blocks Common Crawl bots
# (used by many LLMs)The ideal robots.txt for AI visibility
Here is a complete robots.txt that explicitly allows all important AI crawlers. Copy it and adapt the URLs to your domain:
# robots.txt — optimized for AI visibility # Generated by RadarLLM.ai # Google (standard search) User-agent: Googlebot Allow: / # Google (AI / Gemini / AI Overviews) User-agent: Google-Extended Allow: / # OpenAI (ChatGPT) User-agent: GPTBot Allow: / # OpenAI (ChatGPT Search) User-agent: OAI-SearchBot Allow: / # OpenAI (ChatGPT user browsing) User-agent: ChatGPT-User Allow: / # Anthropic (Claude) User-agent: ClaudeBot Allow: / User-agent: anthropic-ai Allow: / # Perplexity User-agent: PerplexityBot Allow: / # Bing User-agent: Bingbot Allow: / # DuckDuckGo AI User-agent: DuckAssistBot Allow: / # Cohere User-agent: cohere-ai Allow: / # All other bots User-agent: * Allow: / Disallow: /admin/ Disallow: /api/ Disallow: /private/ # Sitemap Sitemap: https://your-site.com/sitemap.xml
Tip: RadarLLM automatically generates an optimized robots.txt for your site during each audit. The file is included in the fix files ZIP.
Common pitfalls
Even if your robots.txt looks correct, other layers can block AI crawlers:
WordPress security plugins
Wordfence, Sucuri and iThemes Security often add rules blocking AI bots in your robots.txt without warning you. Check the "Bot blocking" section of your plugin after every update.
Cloudflare "Bot Fight Mode"
Cloudflare's Bot Fight Mode blocks legitimate AI crawlers by identifying them as threats. Disable it or add exceptions for GPTBot, ClaudeBot and PerplexityBot in your WAF rules.
CDN-level IP blocking
Some CDNs and hosting providers block IP ranges used by AI crawlers. This blocking is invisible in your robots.txt. Check your server logs for blocked requests (403).
Overly broad wildcard rules
A rule "User-agent: * / Disallow: /admin/" seems harmless, but some hosts add "Disallow: /" for all unrecognized agents. Verify that your wildcard is not too restrictive.
What to do after fixing your robots.txt
Fixing your robots.txt is the first step. Here's what comes next:
Wait 4 to 8 weeks
AI crawlers don't come back immediately. After fixing your robots.txt, allow 4 to 8 weeks for GPTBot, ClaudeBot and PerplexityBot to re-index your site.
Monitor with RadarLLM
Run an AI visibility audit to see which LLMs already cite you. Run another audit after 8 weeks to measure progress.
Add structured data
robots.txt opens the door. JSON-LD schemas (FAQPage, Organization) help AIs understand and cite your content. Implement them as a priority.
Create an llms.txt file
robots.txt tells AIs what they can crawl. llms.txt tells them who you are. Both files are complementary and essential in 2026.
Related articles
Is your robots.txt blocking AI?
Scan my website for freeFree scan in under a minute · Fix files with the full audit