Does llms.txt Actually Work? What the 2026 Evidence Says
The short answer is that llms.txt does very little today. Google states the file neither helps nor hurts visibility in its search products, a large-scale server-log study found that the overwhelming majority of published files are never requested by any crawler, and a separate study across hundreds of thousands of domains found no meaningful relationship between having one and being cited in AI answers. This post sets out the actual evidence, because most writing on the subject does not.
What llms.txt is
llms.txt is a proposed convention: a plain-text file at the root of your site giving AI systems a curated map of your most important pages, in Markdown, with short descriptions. The idea is reasonable on its face — models have limited context, your site has navigation and scripts and clutter, so hand them a clean index.
It is a proposal, not a standard. There is no governing body, no compliance requirement, and — this is the crux — no agreement from the systems it addresses to read it.
A related file, ai.txt, is often confused with it. They do different jobs. ai.txt concerns permissions for AI training use; llms.txt concerns content discovery. Adding one does not do the work of the other.
What Google says
Google's position is unusually direct for Google.
Its current documentation tells site owners they do not need llms.txt, Markdown versions of pages, or additional machine-readable files to appear in Search or its generative AI experiences. Google's May 2026 AI optimisation guidance lists llms.txt among tactics site owners can skip, alongside content chunking and AI-specific rewriting. The stated reasoning is that AI Overviews and AI Mode draw from the same index that classic ranking uses, so a separate file changes nothing.
Google goes further than "we don't use it" and says maintaining the file will neither help nor hurt rankings or visibility in Google Search. John Mueller has compared it to the keywords meta tag — a comparison worth sitting with, since the keywords meta tag is the canonical example of a signal publishers maintained for years after search engines stopped reading it.
Source to link in the published post: Google Search Central's AI features and your website guidance — link the current page directly.
VERIFY: Confirm the exact URL and wording on publish day; Google's AI documentation has been revised repeatedly.
What the server logs show
This is the most useful evidence available, because it measures behaviour rather than opinion.
Ahrefs analysed roughly 137,000 domains using bot-analytics data. Around 28% of the domains sampled had a valid llms.txt file. Among those that did, 97% of the files received zero requests during May 2026. Not few requests. Zero. Of the requests that did arrive, AI retrieval crawlers — the systems that actually assemble live AI answers — made up a small fraction.
There is a second observation in the crawler data that is arguably more damning than the first. AI bots do not request llms.txt on domains where it does not exist. A crawler that wanted the file would ask for it and collect 404s. They are not asking. That is the behaviour of systems that have no interest in the file, not systems that check and find it missing.
Source to link: the Ahrefs study.
VERIFY: Link Ahrefs' own publication of the research, not a secondhand summary of it.
What the correlation studies show
SE Ranking examined roughly 300,000 domains and found no statistically significant relationship between having an llms.txt file and how often a domain is cited in AI answers.
That is the question most people actually care about, and the answer is: having the file does not predict getting cited.
Meanwhile adoption is climbing steeply. Originality.ai tracked over three million websites and recorded llms.txt instances rising from around 4,000 to around 36,000 between June 2025 and May 2026 — roughly nine times growth in a year.
Put those two findings together and the picture is clear. The file is being published far faster than it is being read. That gap is the whole story, and it is driven by the marketing industry telling each other this is table stakes.
Sources to link: the SE Ranking study and the Originality.ai tracking data, both directly.
A note on the adoption figures. You will see different percentages quoted — around 28% in one sample, low single digits in another. These are different samples measuring different populations, not contradictory findings. Do not average them or quote one as the adoption rate.
The counter-evidence, stated fairly
The honest version of this post has to include the arguments on the other side, because they exist and are usually omitted by both camps.
- Anthropic recommends llms.txt in its guidance on writing documentation for AI agents.
- OpenAI maintains llms.txt files for its Agents SDK and its agentic commerce work.
- Chrome's Lighthouse added an agentic browsing audit in 2026 that checks for llms.txt, describing it as an emerging convention for giving agents a machine-readable map.
That last one is a genuine internal tension at Google: Search says the file is irrelevant while Chrome's tooling checks for it.
But be precise about what these three facts establish. AI companies publishing llms.txt for their own documentation is not evidence that their models read your llms.txt. Those are different claims, and the first is routinely presented as proof of the second. A developer-tools company publishing a clean index of its own docs is doing something sensible and unrelated to whether a retrieval system parses arbitrary third-party files.
The Chrome audit is more interesting, because it points at agents rather than search. If AI agents that browse on a user's behalf become a meaningful channel — and that is a real possibility rather than a certainty — a machine-readable site map could matter for that use case specifically. That is a forward-looking argument, and it should be labelled as one.
Why it did not work
Worth understanding, because it explains why this is unlikely to reverse on its own.
Standards work when the consuming side adopts them. robots.txt functions because crawlers agreed to read it — the agreement came first, the file second. llms.txt shipped the publisher's half of a handshake that no model provider ever returned.
There is also a deeper problem. Language models build answers from search indexes and crawled pages. They already have your content. A file in which you describe your own site is a lower-trust source than the pages themselves — it is self-reported. Nothing about it gives a retrieval system information it lacks or a reason to weight it more heavily.
The file felt proactive. That is not the same as being useful.
So should you publish one?
The case against spending time on it
If llms.txt is a line item on your SEO roadmap with hours attached, take it off. The evidence does not support the effort. If your platform makes publishing a root file awkward, fighting your host over it is a poor trade for something with no demonstrated return.
More importantly: if anyone has sold you llms.txt as an AI visibility service, or implied that competitors are gaining an advantage through it, that is not supported by anything published. Ask them for the evidence. The studies above are the evidence, and they point the other way.
The narrow case for publishing one anyway
Two honest reasons remain.
It costs almost nothing. If your site is a static build or a CMS where adding a root file takes ten minutes, the downside is ten minutes. Google has explicitly said it will not hurt you.
The writing is the actual value. Producing an llms.txt forces you to state, in a few hundred words, what your organisation is, what it does, what your key pages are and why each matters. Most businesses have never written that down anywhere. The clarity that exercise produces — a clean entity description, consistent terminology, a structured content index — is exactly what does correlate with being cited, because it feeds into how you write everything else. The file is a by-product; the thinking is the asset.
That is a real reason. It is not the reason the file is usually sold.
What we would say to a client: publish it if it takes ten minutes, do the writing exercise properly, and do not count it as AI visibility work. Then spend the time you saved on the list below.
What to do instead, in priority order
Everything here has better evidence behind it than llms.txt does.
- Fix crawlability. If a crawler cannot reach, render and index your pages, no AI system can cite them. This is not glamorous and it is where most sites actually lose.
- Name your authors. Real people, real credentials, bylines on everything, Person schema. Entity credibility is a repeated finding in what generative engines favour, and most Indian business sites name nobody.
- State your entity clearly. One consistent description of what your organisation is, used identically across your site, your Google Business Profile and your social profiles.
- Answer questions completely. Content that resolves a question rather than teasing a click. Question-phrased headings with self-contained 40–70 word answers underneath.
- Publish specifics. Numbers, original data, named examples. Generative models cite specifics and skip generalities.
- Cite external sources. Pages that reference primary sources are treated as more credible than pages that assert.
- Implement real structured data. FAQPage, Article, Organization, Product — validated, not just present.
- Keep it current. Recency is a genuine tiebreaker between credible sources.
If you did all eight and then published an llms.txt as the ninth thing, fine. Doing the ninth thing first is the mistake.
What would change this answer
Stating this explicitly, because a post like this ages:
- A major AI provider confirming its retrieval systems parse third-party llms.txt files
- Crawler logs showing AI retrieval bots requesting the file at scale, including on domains where it does not exist
- A correlation study finding a relationship between publishing the file and citation frequency
- Browsing agents becoming a significant traffic channel and demonstrably using the file
None of those has happened. Any of them would change the recommendation, and this post is reviewed quarterly for exactly that reason.
Frequently asked questions
The evidence says barely. Google states the file neither helps nor hurts visibility in its search products, an analysis of around 137,000 domains found 97% of published files received zero requests in a month, and a study across roughly 300,000 domains found no significant correlation between having the file and being cited in AI answers.
No. Google's documentation states that site owners do not need llms.txt or other AI-specific files to appear in Search or its generative AI features, and that maintaining one will neither help nor hurt rankings. Google's AI optimisation guidance lists it among tactics that can be skipped.
Only if it takes minutes rather than hours, and only if you do not count it as AI visibility work. The genuine benefit is the exercise of writing a clear description of your organisation and its key pages — clarity that helps wherever it is applied. The file itself has no demonstrated effect.
Publishing a clean index of your own documentation for developers and agents is sensible and unrelated to whether a model reads third-party files. These are different claims, and the first is often presented as evidence for the second. No major provider has confirmed that its retrieval systems parse arbitrary sites' llms.txt files.
No. robots.txt controls crawler access and is genuinely honoured by crawlers. ai.txt concerns permissions around AI training use. llms.txt is a proposed content-discovery convention. They serve different purposes and adding one does not accomplish what another does.
Fix crawlability first, then name credentialed authors, state your entity consistently, answer questions completely rather than teasing clicks, publish specific facts and numbers, cite external sources, implement validated structured data, and keep content current. Each of these has better evidence behind it than llms.txt.
Possibly, particularly if AI browsing agents become a significant channel — Chrome's tooling has begun checking for the file in an agentic browsing context, which suggests some expectation of that. That is a forward-looking possibility, not a current return, and it should be described as one.
Ready to build what's next?
Tell us where you're headed. We'll come back with a plan to get there.
Book an intro call