llms.txt: 97% of These Files Went a Whole Month Without a Single Request

Somewhere in the last eighteen months, adding an llms.txt file became one of those tasks that appears on every AI visibility checklist without anyone asking what it does. It is cheap, it takes an afternoon, it looks like doing something. So a lot of sites have one.

In June 2026 Ahrefs published the first large-scale look at whether anything ever asks for the file, using its own crawl and traffic data. The answer was uncomfortable enough that it is worth reading carefully, including the parts that are usually dropped in the retelling.

What was actually measured

The study, by Louise Linehan and Xibeijia Guan, looked at 137,210 domains that received traffic in May 2026. Of those, 38,360 domains, about 28%, published a valid llms.txt. The headline finding is about that second group, not the first: of the roughly 38,000 domains that had published a file, 97% saw no requests for it at all during the month. Only around 1,100 domains got any request whatsoever.

That distinction matters, and it gets mangled constantly. The claim is not “97% of 137,000 sites”. It is “97% of the sites that bothered to publish one”. If you are quoting this number in a meeting, quote it that way, because the mangled version overstates the sample by roughly three and a half times and someone will eventually check.

Ahrefs states three limits on its own work, and all three are worth carrying with the number:

  • The panel is not the web. Ahrefs says its Web Analytics customers skew more technical and SEO-aware than the web at large, so the 28% adoption figure should be treated as an upper bound. Adoption across the actual web is lower.
  • Validity was not tested against the spec. The study did not check whether each file was well formed, only that one existed.
  • Fetched does not mean read. Ahrefs says this outright: many bots may have fetched the file without ever acting on what is inside it. So even the 3% that got requests is an optimistic ceiling on actual use.

Add one more limit that follows from the design: this is a single month, from one vendor’s customer base. It is the best evidence available and it is not a census.

Who is actually asking for it

The more interesting half of the study is the breakdown of who made those requests. Almost all of them, 96%, came from bots. Here is the shape of it.

Requester categoryShare of requestsWhat that actually is
SEO audit tools21.7%Your own crawler, and your competitors’ crawlers, checking the file exists
Unidentified14.9%No usable user agent
General web crawlers13.1%Indiscriminate fetching of anything at a known path
Technology profilers11.6%Tools cataloguing what a site runs
AI agents10.5%Coding assistants pulling documentation into a working session
GEO and AEO tools5.8%The AI visibility tooling industry auditing itself
AI training crawlers5.3%Corpus collection
AI assistants2.5%Live assistant sessions
AI retrieval bots1.1%The systems that actually assemble cited answers

Read the last row twice. The category of bot that decides whether your brand appears inside a generated answer accounts for roughly one percent of requests to an already tiny pool. The named user agents at the top of the list are GPTBot at 4.51% and Claude-Code, which is a coding assistant, not a search product.

The honest conclusion is not “nothing reads llms.txt”. It is sharper than that and more useful: the things that read it are not the things most publishers think they are optimising for. Coding agents read it. Audit tools read it. The retrieval layer behind AI answers almost never does.

Google’s position, in Google’s own words

Google has now written this down explicitly, which removes the room for interpretation. Its guide to optimising for generative AI features has a mythbusting section that says you do not need to create new machine readable files, AI text files, markup or Markdown to appear in Google Search, because Google Search itself does not use them. It then adds the sentence that settles the argument: keeping such files will neither harm nor help your visibility or rankings in Google Search, because Google Search ignores them.

The dating is worth getting right, because it circulates wrongly. The llms.txt paragraph was added to that guide on 15 June 2026, and you can verify that from Google’s own Search Central changelog, which carries a dated entry titled “Clarifying guidance on llms.txt files”. The guide as a whole was published earlier in 2026 and last updated in July; the page itself shows only the last-updated date, so we are not going to assert a publication date we cannot see on the page.

Google had said something similar informally much earlier. In April 2025 John Mueller compared llms.txt to the keywords meta tag, on the grounds that it is a claim by a site owner about its own site which any consumer would have to verify against the site anyway, and that comparison has followed the format around ever since. Note the date on that one: it is a 2025 remark, fourteen months before the crawl data existed, and it should be quoted as a position rather than as evidence.

Where it sits when you weigh the evidence

Cyrus Shepard’s May 2026 review of AI citation ranking factors scored 23 factors by how consistently the available research supports them. llms.txt scored 2, dead last, and the scale in that piece runs from 9.5 down to 2, so it is the observed floor rather than a mark out of ten. His note on it is careful in a way the reposts are not: he is not certain the underlying studies even tested for llms.txt, and he could find no credible evidence or experiment showing it influences citations either way.

We went through the whole scored list, with what each factor costs to fix and who owns it, in the ranked breakdown of all 23 factors.

That is an evidence gap, not a disproof. It is entirely possible that llms.txt does something nobody has measured. It is just not a basis for spending a sprint on it.

What the file was designed for in the first place

Most of the confusion here comes from a category error. Read the original proposal and the motivation is stated plainly: context windows are too small to swallow an entire website, and HTML pages are full of navigation, scripts and chrome that waste those tokens. The file is a curated index that an agent can load at inference time.

That is a developer-tooling problem, not a search problem. And it is exactly what the request data shows: coding assistants and documentation consumers fetch it, because that is the use case it was built for. Which is why the most visible adopters are documentation sites. Anthropic, OpenAI, Stripe, Vercel and Cloudflare all publish one for their developer docs. Worth noting that several of those sites run on the same documentation platform, which generates the file automatically, so some of that adoption is a vendor default rather than a considered endorsement. Cloudflare went further and announced auto-generated llms.txt for customer domains in September 2025, though that shipped as a private beta and we have found no confirmation it reached general availability.

A small detail that says a lot: Ahrefs, which ran the study, does not publish an llms.txt of its own.

These files rot, and ours is proof

The failure mode nobody plans for is staleness. A manually written index of a site diverges from the site within months, and nothing complains. Technical writer Dachary Carey built a small tool to check llms.txt entries against sitemaps and ran it against ten documentation sites. Four covered their documentation completely. Three failed outright. Stripe’s file listed 468 of 3,037 documentation pages, about 15% coverage; Supabase came in around 22%. Ten sites is a probe rather than a study, and the direction is still instructive.

Since this article is about honesty with data, here is ours. As of this writing, dextora.agency serves an llms.txt at the root. It confidently describes an agency offering branding, performance advertising, Meta Ads creatives, Google Ads creatives and conversion optimisation. We do none of those things. It also fails to mention two services we do run, Telegram bots and website migrations, and it does not point at a single article on this site.

Nobody sabotaged it. It was generated once, from a description of the business that was already loose, and then the site changed and the file did not. That is the ordinary lifecycle of this format, and it is a decent illustration of why a self-declared index is weaker evidence than the site itself. A machine that reads our llms.txt learns things about us that are not true. A machine that reads our service pages does not.

So should you keep one

Your situationVerdictWhy
You publish developer documentation or an APIYes, and generate it from the buildThis is the actual use case, and coding agents do fetch it
You run a business site and already have oneFix it or delete itAn inaccurate file is worse than none, because it is confidently wrong
You run a business site and do not have oneNot a priorityThe retrieval systems you care about barely request it
Someone is selling you llms.txt as an AI visibility serviceAsk what the measurement plan isThere is no published evidence of an effect to measure
You can generate and validate it automaticallyFine, it costs nothingAutomation removes the rot problem, which is the only real risk

What to do instead with the same afternoon

The work that has evidence behind it is duller and older. Make sure the page can be read at all without JavaScript, which is a real and common failure and takes ten minutes to check: our curl-based visibility check walks through it. Mark up what the page is with structured data that describes real entities, which scores considerably higher on the same evidence table than llms.txt does. And write pages that answer a question in a self-contained paragraph, which is the mechanic covered in our piece on building a site for AI search.

None of that is new, which is precisely why it is uncomfortable. The appeal of llms.txt was that it looked like a shortcut around the boring work.

The short version

Of roughly 38,000 sites that published an llms.txt, 97% received no request for it in May 2026. The requests that do arrive come overwhelmingly from audit tools and coding assistants; the retrieval bots behind AI answers account for about 1%. Google states plainly that it ignores the file, and neither rewards nor punishes it. The format was designed to fit documentation into a context window, and for that it works. As an AI visibility tactic it has no published evidence behind it, and a stale one actively misinforms. If you keep it, generate it. If you cannot generate it, the site itself is the better index, which is where the effort belongs when we build a corporate site that is meant to be read by machines as well as people.