If your content is marketing, the pages that sell a product or service, you usually shouldn’t block GPTBot, and you should almost never block the crawlers that fetch pages for live AI answers. Blocking makes sense mainly when your content is the product itself: publishing, paid research, proprietary data. The trick is treating training and answer-time crawling as separate decisions, because robots.txt lets you control them separately.
The AI crawlers showing up in your server logs
Most public sites get regular visits from a long list of AI-related user agents. They fall into a few groups, and the names tell you less than you’d hope about what each one does.
| User agent | Operator | What it does | Type |
|---|---|---|---|
| GPTBot | OpenAI | Collects content that may be used to train models | Training |
| ChatGPT-User | OpenAI | Loads a page when a user’s request needs it | User-triggered |
| ClaudeBot | Anthropic | Collects content for model training | Training |
| PerplexityBot | Perplexity | Indexes pages for Perplexity answers; Perplexity says it isn’t used for foundation model training | Search |
| Google-Extended | A robots.txt token, not a separate crawler; governs whether your content helps improve Gemini and its models | Training control | |
| Applebot-Extended | Apple | A token that governs whether Applebot’s crawl trains Apple’s models | Training control |
| CCBot | Common Crawl | Builds an open web archive many AI training datasets draw on | Training, indirect |
| Amazonbot | Amazon | Crawls for Amazon services, which can include AI | Mixed |
Two engines don’t appear here because they don’t have a separate AI crawler. Google’s AI Overviews run on the Googlebot index, and Microsoft Copilot runs on Bing’s. You can’t stay in those search results and out of their AI answers with robots.txt alone.
Operators rename and add agents regularly, so check each one’s crawler documentation before you edit rules.
Training crawlers vs search and user-triggered crawlers
The three types trade very different things.
- Training crawlers feed future models. The payoff for you is slow and diffuse: a model that learned about your company may mention you without searching. There’s no link and no attribution.
- Search crawlers build the index an assistant retrieves from when answering. The payoff is direct: citations, links and inclusion in live answers. Blocking one makes you invisible in that assistant’s search results.
- User-triggered fetchers load a specific URL because a person asked, for example by pasting your pricing page and asking the assistant to compare it with a competitor’s. Blocking these breaks the assistant for someone who is actively evaluating you. Whether robots.txt applies to these requests is also less settled, since a human initiated them.
So the real question isn’t “should I block AI?” It’s two questions: do you want your content in future models, and do you want to show up in answers people read today? For most businesses the second answer is obviously yes, and it drives almost all of the commercial value.
Publishers vs brands: who should block and who should not
Ask one question: is your content the product you sell, or the marketing that sells it?
| Business | Content role | Training crawlers | Search and user crawlers |
|---|---|---|---|
| B2B SaaS, agencies, service firms | Marketing | Allow | Allow |
| Ecommerce and DTC | Marketing | Allow | Allow |
| Ad-supported publisher or niche info site | Product | Blocking is defensible | Usually allow; citations may be the traffic that’s left |
| Paid research, courses, data providers | Product | Block premium paths | Allow on marketing pages |
| Brand with licensed or premium assets | Mixed | Block specific directories | Allow |
If you’re a brand, you want every AI system to know what you sell, who it’s for and roughly what it costs. Blocking training crawlers doesn’t stop models from learning about you. They still learn from review sites, forums and competitors’ comparison pages. You just remove your own version of the story from the mix. When I audit sites for AI visibility, the blocking I find is rarely a deliberate choice. More often it’s a template someone copied or a CDN rule nobody remembers adding.
If you’re a publisher, the value exchange is weaker. An answer that summarizes your article removes the reason to click, and ad revenue depends on the click. Blocking training is a reasonable stance, sometimes paired with licensing discussions. But look at your AI referral traffic before you block search crawlers too, because blocking them removes the citations that still send readers your way.
Google-Extended deserves a separate look. Google’s documentation frames the token around improving Gemini products and the models behind them, and that scope can change. If Gemini matters to your buyers, read the current wording first, because blocking it may cost more than it protects.
How to block selectively in robots.txt
One rule causes most of the mistakes: a crawler obeys only the most specific group that names it and ignores your User-agent: * group entirely. The moment you add a group for PerplexityBot, it stops following your wildcard disallows for cart, admin or internal search. Repeat those lines inside every named group.
Brand default: allow AI, protect utility paths
User-agent: *
User-agent: ChatGPT-User
User-agent: PerplexityBot
Disallow: /cart/
Disallow: /checkout/
Disallow: /account/
Disallow: /search
Naming the AI agents isn’t required, since they’d fall under the wildcard anyway. It documents the decision and keeps the utility disallows attached to them if someone edits the file later.
Content business: block training on premium paths only
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Applebot-Extended
Disallow: /research/
Disallow: /library/
Disallow: /account/
Your public marketing pages stay available for training and answers, while the directories you sell stay out. Use Disallow: / instead only if the whole site is the product.
A few mechanics to get right:
- robots.txt applies per host.
docs.example.comandblog.example.comeach need their own file. - On Shopify, you edit it through the
robots.txt.liquidtheme template, not by uploading a file. - Your CDN or firewall can block bots before robots.txt is ever read. Check bot management settings, especially any broad rule aimed at AI or “bot” user agents, which may catch search and user fetchers too.
Crawler access is one line in a broader technical SEO audit checklist, but it’s the one that can quietly remove you from AI answers.
Checking server logs for AI bot activity
Your robots.txt says what you intended. Your logs show what’s actually happening.
- Get the raw logs. Pull access logs from your server, CDN or host. Hosted platforms often don’t expose raw logs, so CDN bot analytics fills the gap.
- Count hits by agent. A quick pass on a standard access log:
grep -oiE "GPTBot|ChatGPT-User|ClaudeBot|PerplexityBot|CCBot|Amazonbot" access.log | sort | uniq -c | sort -rn
- Verify identity. User agent strings are trivial to fake. Major operators publish IP ranges or support reverse DNS lookups, so treat hits from outside those as unverified.
- Read the status codes. A 403 or 429 on a bot you meant to allow means your firewall and robots.txt disagree. A 200 on a path you disallowed means the rule is wrong or the bot isn’t complying.
- Look at which URLs get fetched. User-triggered fetches on pricing, comparison and documentation pages mean real people are asking assistants about you. That’s a strong reason not to block.
- Watch the volume. If one crawler hammers your server, rate-limit it at the CDN rather than blocking everything with “bot” in its name.
Check again a week or two after any change, then quarterly, because new agents appear.
What robots.txt does and does not prevent
robots.txt is a public request, not a lock. It tells compliant crawlers which paths to skip from now on, and it lets you treat each operator’s training and search bots differently. That’s all it does.
It does not:
- Remove content that was already collected or models that were already trained
- Cover copies of your content on other sites, such as syndication partners, scrapers or quotes
- Stop a model from describing your brand using everything else written about you
- Reliably bind user-triggered fetchers, which may not follow it
- Keep you out of AI Overviews, which follow Google Search’s own indexing and snippet controls
- Stop crawlers that choose to ignore it
For content that truly has to stay out, use authentication, firewall rules matched to verified bot IPs, or licensing terms. Other text files pitched at AI tools don’t replace it; the guide to AI site files covers what they do and don’t do.
A sensible default for most businesses
For a SaaS company, an ecommerce brand or a service business, this is where I’d start:
- Allow Googlebot, Bingbot and Applebot for search, AI Overviews, Copilot and Siri
- Allow search crawlers that power AI answers, such as PerplexityBot
- Allow user-triggered fetchers such as ChatGPT-User
- Decide on training crawlers with one question: is this content the product or the marketing?
- If you block training, block by directory, not site-wide
- Disallow cart, checkout, account, internal search and staging paths in every group
- Match CDN and firewall settings to robots.txt
- Confirm in logs after each change and review quarterly for new agents
- Record the decision and its owner so the next redesign or migration doesn’t undo it
Crawler access is the first thing I check in GEO and AI search visibility work, because no amount of content or mention building helps if the assistants can’t read the page.
Get it built
If you’re not sure what your robots.txt, CDN and firewall are telling AI crawlers, the Growth Audit checks crawler access alongside the rest of your search and AI visibility. It’s $1,500 fixed and credited if we continue. See pricing or get in touch.