Skip to content
Can Elmas

GEO · 8 min read

Should You Block GPTBot and Other AI Crawlers? A Decision Guide

TL;DR

If your content is marketing, don't block the crawlers that fetch pages for AI answers, such as PerplexityBot and ChatGPT-User, and think twice before blocking training crawlers like GPTBot. Blocking training makes sense mainly when your content is the product you sell. Write separate robots.txt rules per bot, align your CDN settings, then verify in server logs.

· Published · Updated

If your content is marketing, the pages that sell a product or service, you usually shouldn’t block GPTBot, and you should almost never block the crawlers that fetch pages for live AI answers. Blocking makes sense mainly when your content is the product itself: publishing, paid research, proprietary data. The trick is treating training and answer-time crawling as separate decisions, because robots.txt lets you control them separately.

The AI crawlers showing up in your server logs

Most public sites get regular visits from a long list of AI-related user agents. They fall into a few groups, and the names tell you less than you’d hope about what each one does.

User agentOperatorWhat it doesType
GPTBotOpenAICollects content that may be used to train modelsTraining
ChatGPT-UserOpenAILoads a page when a user’s request needs itUser-triggered
ClaudeBotAnthropicCollects content for model trainingTraining
PerplexityBotPerplexityIndexes pages for Perplexity answers; Perplexity says it isn’t used for foundation model trainingSearch
Google-ExtendedGoogleA robots.txt token, not a separate crawler; governs whether your content helps improve Gemini and its modelsTraining control
Applebot-ExtendedAppleA token that governs whether Applebot’s crawl trains Apple’s modelsTraining control
CCBotCommon CrawlBuilds an open web archive many AI training datasets draw onTraining, indirect
AmazonbotAmazonCrawls for Amazon services, which can include AIMixed

Two engines don’t appear here because they don’t have a separate AI crawler. Google’s AI Overviews run on the Googlebot index, and Microsoft Copilot runs on Bing’s. You can’t stay in those search results and out of their AI answers with robots.txt alone.

Operators rename and add agents regularly, so check each one’s crawler documentation before you edit rules.

Training crawlers vs search and user-triggered crawlers

The three types trade very different things.

  • Training crawlers feed future models. The payoff for you is slow and diffuse: a model that learned about your company may mention you without searching. There’s no link and no attribution.
  • Search crawlers build the index an assistant retrieves from when answering. The payoff is direct: citations, links and inclusion in live answers. Blocking one makes you invisible in that assistant’s search results.
  • User-triggered fetchers load a specific URL because a person asked, for example by pasting your pricing page and asking the assistant to compare it with a competitor’s. Blocking these breaks the assistant for someone who is actively evaluating you. Whether robots.txt applies to these requests is also less settled, since a human initiated them.

So the real question isn’t “should I block AI?” It’s two questions: do you want your content in future models, and do you want to show up in answers people read today? For most businesses the second answer is obviously yes, and it drives almost all of the commercial value.

Publishers vs brands: who should block and who should not

Ask one question: is your content the product you sell, or the marketing that sells it?

BusinessContent roleTraining crawlersSearch and user crawlers
B2B SaaS, agencies, service firmsMarketingAllowAllow
Ecommerce and DTCMarketingAllowAllow
Ad-supported publisher or niche info siteProductBlocking is defensibleUsually allow; citations may be the traffic that’s left
Paid research, courses, data providersProductBlock premium pathsAllow on marketing pages
Brand with licensed or premium assetsMixedBlock specific directoriesAllow

If you’re a brand, you want every AI system to know what you sell, who it’s for and roughly what it costs. Blocking training crawlers doesn’t stop models from learning about you. They still learn from review sites, forums and competitors’ comparison pages. You just remove your own version of the story from the mix. When I audit sites for AI visibility, the blocking I find is rarely a deliberate choice. More often it’s a template someone copied or a CDN rule nobody remembers adding.

If you’re a publisher, the value exchange is weaker. An answer that summarizes your article removes the reason to click, and ad revenue depends on the click. Blocking training is a reasonable stance, sometimes paired with licensing discussions. But look at your AI referral traffic before you block search crawlers too, because blocking them removes the citations that still send readers your way.

Google-Extended deserves a separate look. Google’s documentation frames the token around improving Gemini products and the models behind them, and that scope can change. If Gemini matters to your buyers, read the current wording first, because blocking it may cost more than it protects.

How to block selectively in robots.txt

One rule causes most of the mistakes: a crawler obeys only the most specific group that names it and ignores your User-agent: * group entirely. The moment you add a group for PerplexityBot, it stops following your wildcard disallows for cart, admin or internal search. Repeat those lines inside every named group.

Brand default: allow AI, protect utility paths

User-agent: *
User-agent: ChatGPT-User
User-agent: PerplexityBot
Disallow: /cart/
Disallow: /checkout/
Disallow: /account/
Disallow: /search

Naming the AI agents isn’t required, since they’d fall under the wildcard anyway. It documents the decision and keeps the utility disallows attached to them if someone edits the file later.

Content business: block training on premium paths only

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Applebot-Extended
Disallow: /research/
Disallow: /library/
Disallow: /account/

Your public marketing pages stay available for training and answers, while the directories you sell stay out. Use Disallow: / instead only if the whole site is the product.

A few mechanics to get right:

  • robots.txt applies per host. docs.example.com and blog.example.com each need their own file.
  • On Shopify, you edit it through the robots.txt.liquid theme template, not by uploading a file.
  • Your CDN or firewall can block bots before robots.txt is ever read. Check bot management settings, especially any broad rule aimed at AI or “bot” user agents, which may catch search and user fetchers too.

Crawler access is one line in a broader technical SEO audit checklist, but it’s the one that can quietly remove you from AI answers.

Checking server logs for AI bot activity

Your robots.txt says what you intended. Your logs show what’s actually happening.

  1. Get the raw logs. Pull access logs from your server, CDN or host. Hosted platforms often don’t expose raw logs, so CDN bot analytics fills the gap.
  2. Count hits by agent. A quick pass on a standard access log:
grep -oiE "GPTBot|ChatGPT-User|ClaudeBot|PerplexityBot|CCBot|Amazonbot" access.log | sort | uniq -c | sort -rn
  1. Verify identity. User agent strings are trivial to fake. Major operators publish IP ranges or support reverse DNS lookups, so treat hits from outside those as unverified.
  2. Read the status codes. A 403 or 429 on a bot you meant to allow means your firewall and robots.txt disagree. A 200 on a path you disallowed means the rule is wrong or the bot isn’t complying.
  3. Look at which URLs get fetched. User-triggered fetches on pricing, comparison and documentation pages mean real people are asking assistants about you. That’s a strong reason not to block.
  4. Watch the volume. If one crawler hammers your server, rate-limit it at the CDN rather than blocking everything with “bot” in its name.

Check again a week or two after any change, then quarterly, because new agents appear.

What robots.txt does and does not prevent

robots.txt is a public request, not a lock. It tells compliant crawlers which paths to skip from now on, and it lets you treat each operator’s training and search bots differently. That’s all it does.

It does not:

  • Remove content that was already collected or models that were already trained
  • Cover copies of your content on other sites, such as syndication partners, scrapers or quotes
  • Stop a model from describing your brand using everything else written about you
  • Reliably bind user-triggered fetchers, which may not follow it
  • Keep you out of AI Overviews, which follow Google Search’s own indexing and snippet controls
  • Stop crawlers that choose to ignore it

For content that truly has to stay out, use authentication, firewall rules matched to verified bot IPs, or licensing terms. Other text files pitched at AI tools don’t replace it; the guide to AI site files covers what they do and don’t do.

A sensible default for most businesses

For a SaaS company, an ecommerce brand or a service business, this is where I’d start:

  • Allow Googlebot, Bingbot and Applebot for search, AI Overviews, Copilot and Siri
  • Allow search crawlers that power AI answers, such as PerplexityBot
  • Allow user-triggered fetchers such as ChatGPT-User
  • Decide on training crawlers with one question: is this content the product or the marketing?
  • If you block training, block by directory, not site-wide
  • Disallow cart, checkout, account, internal search and staging paths in every group
  • Match CDN and firewall settings to robots.txt
  • Confirm in logs after each change and review quarterly for new agents
  • Record the decision and its owner so the next redesign or migration doesn’t undo it

Crawler access is the first thing I check in GEO and AI search visibility work, because no amount of content or mention building helps if the assistants can’t read the page.

Get it built

If you’re not sure what your robots.txt, CDN and firewall are telling AI crawlers, the Growth Audit checks crawler access alongside the rest of your search and AI visibility. It’s $1,500 fixed and credited if we continue. See pricing or get in touch.

FAQ

Frequently Asked Questions

Will blocking GPTBot remove my site from ChatGPT?

Not by itself. GPTBot collects training data, while ChatGPT loads pages for live answers through ChatGPT-User, which has its own robots.txt rule. Block ChatGPT-User and ChatGPT can stop loading your pages when a user's question needs them.

Does blocking Google-Extended keep me out of AI Overviews?

No. AI Overviews draw on the regular Google Search index crawled by Googlebot, and Google-Extended doesn't affect Search. The levers there are Search controls such as nosnippet or noindex, which also limit or remove your normal search listing.

Can AI companies still use content I've blocked in robots.txt?

Yes, in several ways. robots.txt doesn't remove content already collected, doesn't cover copies of your text on other sites, and some user-triggered fetchers may not apply it. Content that must stay out belongs behind a login or a firewall rule.

Should an ecommerce store block AI crawlers?

Usually not. Product pages, prices and reviews are marketing, and an AI assistant can only recommend products it can read. Disallow cart, checkout and account paths for every bot and leave the rest open.

Work with me

Let’s find your biggest growth lever

Tell me about your growth challenge. I’ll tell you honestly if I can help — and if I can’t, who can.

  • ✓ No obligation
  • ✓ No sales script
  • ✓ Honest feedback
  • ✓ Clear next steps