RadarLLM
2026-04-087 minTechnical

Your robots.txt is blocking ChatGPT without you knowing

Cyril Cretin

Cyril Cretin

Founder RadarLLM.ai

Your website may be invisible to ChatGPT, Claude, Perplexity AND Gemini right now. The most common cause: your robots.txt. Many hosting providers, CDNs and security plugins block AI crawlers by default, without telling you. The result: AIs simply cannot read your site.

5.89% of all websites block GPTBot. And among major publishers, 79% block at least one AI bot. If you've never checked your robots.txt, there's a strong chance you're part of these statistics.

AI bots you need to know

Each AI platform uses one or more bots to crawl the web. Here is the complete list of bots your robots.txt must allow:

BotPlatformPurposeAllow?
GPTBotChatGPTTraining + SearchYes
OAI-SearchBotChatGPT SearchReal-time searchYes
ChatGPT-UserChatGPTUser browsingYes
ClaudeBotClaudeTraining + citationsYes
anthropic-aiAnthropicTrainingYes
PerplexityBotPerplexitySearch + citationsYes
Google-ExtendedGeminiTraining + AI OverviewsYes
GooglebotGoogleStandard searchYes

How to check if your robots.txt blocks AI

1.

Open your browser and go to your-site.com/robots.txt

2.

Search for each bot name listed above (GPTBot, ClaudeBot, PerplexityBot, Google-Extended...)

3.

Look for "Disallow: /" rules associated with these bots — that's the blocking signal

Here are the most common patterns that block AI crawlers:

# Examples of rules that BLOCK AI crawlers:

User-agent: GPTBot
Disallow: /          # ← Blocks ChatGPT

User-agent: ClaudeBot
Disallow: /          # ← Blocks Claude

User-agent: *
Disallow: /          # ← Blocks EVERYTHING, including AI

User-agent: CCBot
Disallow: /          # ← Blocks Common Crawl bots
                     #   (used by many LLMs)

The ideal robots.txt for AI visibility

Here is a complete robots.txt that explicitly allows all important AI crawlers. Copy it and adapt the URLs to your domain:

# robots.txt — optimized for AI visibility
# Generated by RadarLLM.ai

# Google (standard search)
User-agent: Googlebot
Allow: /

# Google (AI / Gemini / AI Overviews)
User-agent: Google-Extended
Allow: /

# OpenAI (ChatGPT)
User-agent: GPTBot
Allow: /

# OpenAI (ChatGPT Search)
User-agent: OAI-SearchBot
Allow: /

# OpenAI (ChatGPT user browsing)
User-agent: ChatGPT-User
Allow: /

# Anthropic (Claude)
User-agent: ClaudeBot
Allow: /
User-agent: anthropic-ai
Allow: /

# Perplexity
User-agent: PerplexityBot
Allow: /

# Bing
User-agent: Bingbot
Allow: /

# DuckDuckGo AI
User-agent: DuckAssistBot
Allow: /

# Cohere
User-agent: cohere-ai
Allow: /

# All other bots
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /api/
Disallow: /private/

# Sitemap
Sitemap: https://your-site.com/sitemap.xml

Tip: RadarLLM automatically generates an optimized robots.txt for your site during each audit. The file is included in the fix files ZIP.

Common pitfalls

Even if your robots.txt looks correct, other layers can block AI crawlers:

01

WordPress security plugins

Wordfence, Sucuri and iThemes Security often add rules blocking AI bots in your robots.txt without warning you. Check the "Bot blocking" section of your plugin after every update.

02

Cloudflare "Bot Fight Mode"

Cloudflare's Bot Fight Mode blocks legitimate AI crawlers by identifying them as threats. Disable it or add exceptions for GPTBot, ClaudeBot and PerplexityBot in your WAF rules.

03

CDN-level IP blocking

Some CDNs and hosting providers block IP ranges used by AI crawlers. This blocking is invisible in your robots.txt. Check your server logs for blocked requests (403).

04

Overly broad wildcard rules

A rule "User-agent: * / Disallow: /admin/" seems harmless, but some hosts add "Disallow: /" for all unrecognized agents. Verify that your wildcard is not too restrictive.

What to do after fixing your robots.txt

Fixing your robots.txt is the first step. Here's what comes next:

01

Wait 4 to 8 weeks

AI crawlers don't come back immediately. After fixing your robots.txt, allow 4 to 8 weeks for GPTBot, ClaudeBot and PerplexityBot to re-index your site.

02

Monitor with RadarLLM

Run an AI visibility audit to see which LLMs already cite you. Run another audit after 8 weeks to measure progress.

03

Add structured data

robots.txt opens the door. JSON-LD schemas (FAQPage, Organization) help AIs understand and cite your content. Implement them as a priority.

04

Create an llms.txt file

robots.txt tells AIs what they can crawl. llms.txt tells them who you are. Both files are complementary and essential in 2026.

Is your robots.txt blocking AI?

Scan my website for free

Free scan in under a minute · Fix files with the full audit