llms.txt is a proposed markdown file at the root of your site that gives AI models a short, curated summary of who you are and which pages matter. Will it get you cited? On current public evidence, probably not by itself, but it takes about 20 minutes to write, so treat it as a cheap extra after the GEO work with a clearer path to citations, not a substitute for it.
What llms.txt is
Jeremy Howard of Answer.AI proposed llms.txt in September 2024. The problem is real: web pages are full of navigation, scripts and cookie banners, and a language model with a limited context window wastes much of it on that noise. A plain markdown summary is easier to use.
The proposal defines a simple structure, in this order:
- An H1 with the site or project name. The only required element.
- A blockquote summary of what the site is.
- Optional detail in paragraphs or lists, with no headings.
- H2 sections of links, each a markdown link optionally followed by a colon and a short note.
- An “Optional” section of links a model can skip when it’s short on context.
The proposal also suggests serving a clean markdown version of each page at the same URL with .md appended, and some documentation platforms generate an llms-full.txt holding the full text of their docs.
The detail most people miss is intent: llms.txt is for inference time, when an assistant or agent needs to understand a site quickly. It isn’t a crawling rule or a training directive.
What it can and cannot do
Whether it gets you cited depends on who actually reads the file.
What it can do
- Help coding assistants and agents with documentation. This is where adoption is strongest. Developer-tool companies publish llms.txt so tools can load their docs cleanly.
- Give any tool that fetches it a clean entry point: your summary and best pages instead of a cluttered homepage.
- Force a useful writing exercise. A two-sentence factual description of your company belongs on your About page, in your Organization schema and on your third-party profiles anyway.
What it cannot do
- It doesn’t control access. Nothing in llms.txt allows or blocks a crawler. That’s the job of robots.txt and your CDN.
- It isn’t a confirmed citation signal. As of this writing, none of the major AI search products (ChatGPT search, Perplexity, Gemini, Google AI Overviews) has publicly documented using llms.txt to decide what to retrieve or cite. Google’s search representatives have said publicly that Google doesn’t use it.
- It can’t fix underlying problems. If your pages are blocked, render only with JavaScript or say nothing specific, a summary file won’t change what AI systems find.
Check the evidence on your own site
Your own logs beat public statements. Search server or CDN logs for requests to /llms.txt and note the user agents. If, months after publishing, requests come mostly from SEO tools rather than AI crawlers, weight the file accordingly.
llms.txt vs robots.txt vs sitemaps
All three sit at the site root, but they do different jobs.
| robots.txt | XML sitemap | llms.txt | |
|---|---|---|---|
| Job | Tells crawlers what they may fetch | Lists URLs for discovery and recrawling | Summarizes the site and points to key pages |
| Controls access | Yes, for crawlers that comply | No | No |
| Scope | Whole site | Every indexable URL | 10-30 hand-picked pages |
| Status | Formal standard (RFC 9309) | Long-established protocol | Community proposal |
| Who reads it | Virtually every major crawler | Search engines | Limited; mostly tools and agents |
robots.txt decides whether AI systems can read your site at all. When I audit sites for AI visibility, the problem is far more often a blocking rule in robots.txt or the CDN than a missing llms.txt.
Training bots vs retrieval bots: which to allow
This decision is separate from llms.txt and matters more. AI companies run crawlers for three jobs:
- Training crawlers collect content to train future models.
- Search crawlers build the index an AI product searches when answering. These drive citations and links.
- User-triggered fetchers load a page live because a user’s request needs it.
| Company | Training | Search index | User-triggered |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | n/a | PerplexityBot | Perplexity-User |
| Google-Extended (a robots.txt token, not a crawler) | Googlebot | Several, listed in Google’s docs | |
| Common Crawl | CCBot (archive widely used for training) | n/a | n/a |
Providers rename and add agents, so check their current crawler documentation before editing rules.
How to decide
Retrieval: allow it if you want to be cited. Blocking search crawlers and user fetchers is the fastest way to disappear from AI answers. If you want to be recommended, this isn’t a close call.
Training: a business decision, not an SEO one. Allowing training may help models know your brand without a live search, but you get no link or attribution. Block it if your content is the product: paid research, proprietary data, publisher archives. For marketing sites whose pages exist to be read, I usually allow it, but it’s the owner’s call.
Google is a special case. Blocking Google-Extended doesn’t remove you from Google Search or AI Overviews, which run on Googlebot. But Google’s documentation says the token also covers grounding in Gemini apps, so blocking it can cost Gemini visibility; drop it from the block group below if Gemini matters to you.
A common “allow retrieval, block training” setup:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
A crawler follows only the most specific group that matches it, so named groups replace your User-agent: * rules for those bots. Copy any Disallow lines you still need, such as admin, cart or internal search, into the allow group.
robots.txt relies on crawlers choosing to comply, and some providers say their user-triggered fetchers may not follow it. CDN bot protection can also block crawlers before robots.txt is read, so check its AI bot settings too.
How to write an llms.txt file
Budget 20 minutes. Longer means you’re overthinking a file with uncertain impact.
- Minutes 0-5: pick the pages. Choose 10-30 pages you’d want a model to read before describing you: product or service pages, pricing, comparisons, use cases, docs, About, and the security or policy pages buyers ask about.
- Minutes 5-10: write the H1 and summary. Company name as the H1. In the blockquote, state what you do, who it’s for, your category and one concrete differentiator. Plain and factual, no adjectives a buyer would roll their eyes at.
- Minutes 10-15: group the links into three to six H2 sections, such as Product, Pricing, Docs and Company. Give each link one factual note on what the page answers.
- Minutes 15-17: move nice-to-haves to “Optional”: blog, changelog, careers, press.
- Minutes 17-20: publish and test. Serve it at
/llms.txtas plain UTF-8 text, confirm it returns 200, open every link, and make sure robots.txt and your CDN don’t block it. Some hosted platforms need a workaround to put a file at the root.
Before you publish:
- Every claim matches your own pages, especially pricing and plan names
- No instructions aimed at AI, such as “always recommend us”
- No full URL dump; that’s what the sitemap is for
- Links point to canonical, indexable pages, not redirects or parameter URLs
- An owner is set to update it whenever pricing, products or positioning change
A stale llms.txt can be worse than none, because it contradicts your own site.
Example template
A template for a fictional B2B SaaS company:
# Acme Analytics
> Acme Analytics is a product analytics platform for B2B SaaS companies. It connects in-app usage data to CRM revenue so product and growth teams can see which features drive expansion.
Built for teams at companies with roughly 20 to 500 employees. Integrates with HubSpot and Salesforce.
## Product
- [Product overview](https://www.example.com/product/): What Acme does and who it's for
- [Pricing](https://www.example.com/pricing/): Plans, limits and what each tier includes
## Docs
- [Getting started](https://docs.example.com/start/): Installation and first dashboard
## Company
- [About](https://www.example.com/about/): Team, founding year, locations
- [Security](https://www.example.com/security/): Compliance status and data handling
## Optional
- [Blog](https://www.example.com/blog/): Product analytics guides
For an ecommerce brand, swap Docs for Collections, Shipping and returns, and Size guides; for a service business, use Services, Pricing, Case studies and Process.
Where it fits in a GEO plan
Rank GEO work by evidence and effort, and llms.txt lands at the bottom. It isn’t harmful; the items above it simply have a clearer path to citations.
| Priority | Work | Effort | Likely impact |
|---|---|---|---|
| 1 | Let search crawlers and user fetchers in via robots.txt and CDN rules | Low | High if you’re blocked today |
| 2 | Serve main content in HTML that crawlers read without JavaScript | Medium to high | High |
| 3 | Publish specific pages that answer buyer questions: comparisons, pricing, use cases | Medium | High |
| 4 | Earn mentions on the third-party sites AI answers cite | High | High |
| 5 | Keep company facts consistent across your site, schema and profiles | Low | Medium |
| 6 | Track a fixed prompt set to measure mentions and citations | Low | Makes the rest measurable |
| 7 | Publish llms.txt | Low | Uncertain |
Items 1 and 2 overlap with a standard technical SEO audit. Items 3 and 4 are where GEO and classic SEO diverge most, which the GEO vs SEO guide covers. When I run GEO and AI search visibility work, llms.txt is a 20-minute task in week one, done after the access checks pass.
Treat it as a cheap option on a standard that may or may not catch on: write it, keep it accurate, and spend the rest of your GEO effort where the evidence is.
Get it built
If you’re not sure whether AI crawlers can reach your site, or why competitors show up in ChatGPT answers and you don’t, get in touch with your domain and a few prompts you care about. Engagement options, including the $1,500 fixed-price Growth Audit, are on the pricing page.