Does llms.txt Do Anything? I Checked the Logs
Both sides of this argument have been running on first principles for two years. Three server log studies, one correlation study and one live audit later, the argument is over. Here is the evidence, including the part that cuts against me.
- Across 137,210 domains in May 2026, AI retrieval bots sent 233 requests to llms.txt files. Bots that exist to audit llms.txt sent 6,847. The audit industry reads the file 29 times more often than the engines do.
- Three independent server log studies (Ahrefs, Adobe Experience Manager customer logs, and a 90 day single site test) all land in the same place, and two of them independently report 1.1% of llms.txt traffic coming from verifiable AI agents.
- The adoption numbers in circulation disagree by 13x, from 28% down to 2.13%, because the studies count different things and HTTP 200 is not proof a file exists.
- I checked 28 domains by hand on 2026-07-28. robots.txt returned 200 on all 28. A real llms.txt returned on 12. OpenAI, Perplexity, Google and anthropic.com all 404.
- The one surviving use case is developer documentation read by coding agents, which is roughly 10x better evidenced than the search case and is still not a citation lever.
The number nobody has published
Ahrefs Bot Analytics logged roughly 22,000 requests to llms.txt files across 137,210 domains in May 2026, then sorted every request by what sent it. Twelve categories came back. AI retrieval bots, the crawlers that fetch a page to answer a live user question, finished last with 233 requests.
Now add the categories that exist purely to check whether the file is there. SEO audit tools sent 4,776 requests. GEO and AEO readiness scanners sent 1,278. Bots whose stated job is to scan, validate or catalogue llms.txt files sent 793. That is 6,847 requests.
The audit industry reads your llms.txt file about 29 times more often than the engines do. That subtraction is sitting in a public table and I have not seen anyone run it. It is the most honest one line description of what this file currently is: a compliance artifact, mostly read by the tools that told you to install it.
I installed it. On two client sites, in early 2025, on exactly the reasoning everyone used. The logic was fine. The logs were not.
llms.txt does not work as an AI search visibility lever. Three independent server log studies and one 300,000 domain correlation study all point the same way, and no major engine has committed to reading it. The one use that survives the evidence is developer documentation consumed by coding agents, which is a different job from getting cited in an answer.
What the file was actually proposed to do
Jeremy Howard published the llms.txt proposal on 3 September 2024. Read it and you find something more modest than what got sold on top of it.
Our expectation is that llms.txt will mainly be useful for inference, i.e. at the time a user is seeking assistance.
The stated problem is context windows, not ranking. The proposal frames the file as a curated markdown map so a model can use a site without ingesting all of it. Nothing in it claims a citation benefit, and nothing in it claims any engine consumes it. The proposal never overpromised. The industry that resold it did.
That distinction matters when you evaluate the counterargument you will hear most: that Anthropic, Stripe, Vercel and Cloudflare all publish one, so it must be working. They publish one because Mintlify shipped it as a default across the documentation sites it hosts. Adoption in that cohort is a hosting decision, not a marketing result. The same pattern shows up in the general web numbers, where the HTTP Archive Web Almanac found 39.6% of valid llms.txt files were generated by the All in One SEO plugin. Most llms.txt files on the internet were not chosen. They arrived.
Five datasets, and they do not disagree
This is the part the debate has been missing. There is not one study. There are five, using four different methods, and they converge.
| Study | Method | Scale | Headline finding |
|---|---|---|---|
| Ahrefs, Jun 2026 | Bot Analytics logs | 137,210 domains, May 2026 | 97% of llms.txt files got zero requests. AI retrieval bots were 1.1% of traffic to the file |
| Flavio Longato, Jun 2026 | Adobe Experience Manager customer server logs | 22,494 requests, 30 day window | Verifiable LLM agents were 258 requests, 1.1% of all traffic to the file |
| Otterly.ai | Single site server logs, 90 days | 62,100+ AI bot hits | 84 hits to /llms.txt, about 0.1% of AI bot visits, against an average page at roughly 265 |
| SEO Depths, Simone De Palma | PHP server logs, two sites | Small site plus a 100K+ URL travel site | 10 requests in 30 days on the large site. None measurable on the small one |
| SE Ranking, Yulia Deda | Spearman correlation plus XGBoost with SHAP | ~300,000 domains | No correlation between llms.txt presence and AI citation frequency |
The convergence worth staring at is the 1.1%. Ahrefs measured it against a commercial bot analytics panel. Longato measured it against enterprise CMS server logs at Adobe. Different populations, different classifiers, different months, same figure. When two unrelated methods land on the same number, the number is usually real. That is a stronger evidentiary position than most claims in generative engine optimization ever reach, and it happens to be pointing at a negative result.
Longato also found the detail that kills the mechanism outright, not just the volume.
I did not find a single request anywhere in the server logs whose referrer was a /llms.txt URL.
The whole premise is that a model reads the file and then fetches the curated URLs inside it. If that were happening at any scale, there would be a referrer trail. There is not one. Crawlers enter at the homepage and follow links, the way crawlers have always worked, which is also why crawler access and internal linking still outrank every novelty file on the priority list.
The adoption numbers are broken, and that is the real story
Here is where I part company with everyone writing about this, including the people I agree with. The volume debate is settled. The adoption debate is a mess, and the mess is being repeated as fact.
That is a 13x spread on a simple question: does this file exist. Longato's independent probe of 4,685 hosts came back at 2.9% returning a working 200 response, which corroborates HTTP Archive and not Ahrefs. Then he tightened the test. Requiring plain text dropped it to 111 hosts. Requiring actual content inside the file dropped it to 20. Twenty out of 4,685.
Two things explain the gap, and only one of them is usually mentioned.
The first is sampling. Ahrefs measured 137,210 domains inside Ahrefs Web Analytics. Those are sites run by people who buy SEO software. HTTP Archive crawls the web. A panel of SEO tool customers adopting a novelty SEO file at 13x the rate of the open web is exactly what you would predict. That is my inference, not Ahrefs' claim, and I would hold it loosely. But anyone quoting "28% of sites have llms.txt" as a fact about the web is quoting a fact about a customer base.
The second is that HTTP 200 does not mean the file exists, and I can show you that on a real domain in the next section. Any adoption study that counts status codes without checking content type and body is counting soft 404s as adoption. This is the same class of measurement error that makes AI visibility sample sizes so easy to get wrong, and it is why I now check the bytes rather than the code.
I checked 28 domains by hand
On 28 July 2026 I requested /llms.txt, /robots.txt and /sitemap.xml from 28 domains with curl, recording status code, content type and response size for each. The domain set was deliberately biased toward exactly the companies most likely to have one: nine model and engine vendors, eight SEO software vendors, four developer documentation hosts, and seven large general web properties. This is a convenience sample of 28. It proves nothing about the web and it was not meant to. It was meant to answer one narrow question that a log study cannot: who bothered.
The control worked. robots.txt returned HTTP 200 on 28 of 28. The XML sitemap returned 200 at the root path on 23 of 28. So the method reaches these hosts fine.
A real llms.txt, meaning HTTP 200 with a text content type and actual markdown inside, came back on 12 of 28. Thirteen returned HTTP 200. The thirteenth was reddit.com, which serves an HTML error page under a 200 status code. A naive audit script counts that as adoption. That single case is the measurement bug from the previous section, live.
The distribution is the interesting part.
| Group | Domains | Served a real llms.txt |
|---|---|---|
| Model and engine vendors | 9 | 2 (claude.com, mistral.ai) |
| Developer documentation hosts | 4 | 4 |
| SEO software vendors | 8 | 3 |
| Large general web properties | 7 | 3 |
openai.com, perplexity.ai, google.com, microsoft.com, bing.com, x.ai and anthropic.com itself all returned 404. So did developers.google.com, the domain that hosts Google's own guidance on optimizing for generative AI. OpenAI's 404 for /llms.txt shipped 148,942 bytes of HTML to tell me the file was not there.
Run the same three requests against your own domain before you accept any vendor's audit of it. The command is at the bottom of this page.
What the engines have actually said
Google put it in the Search Central documentation changelog on 15 June 2026, under the heading "Clarifying guidance on llms.txt files":
> while these files aren't needed for Google Search (and won't negatively or positively impact your visibility or rankings), it's fine if you want to maintain these files for other services or systems that use them.
That is a permission slip, not an endorsement. Read alongside Google's AI features documentation, which states there are no additional requirements to appear in AI Overviews or AI Mode, the position is consistent: the AI surfaces run on the ordinary index, and the only hard prerequisite Google names is that a page be indexed and eligible to show with a snippet.
Google's John Mueller, as reported by PPC Land, made the same argument from the same evidence I am using here, which is the part I find persuasive.
none of the AI services have said they're using LLMs.TXT (and you can tell when you look at your server logs that they don't even check for it)
And the position of the company that ran the largest study:
If your goal is showing up in ChatGPT, Perplexity, or AI Overviews, an llms.txt file is largely decoration.
No engine has ever announced support. OpenAI has not. Anthropic has not. Google has explicitly said it does not use it. Absence of an announcement is weak evidence on its own. Absence of an announcement plus three log studies showing nobody fetches the file is not weak at all.
If someone quoted you for an llms.txt build, ask them for the fetch counts. Or bring me the domain and I will run the access and citation audit that actually moves numbers.
Book a call→The steelman: where llms.txt genuinely earns its place
I would rather publish the case against my own position than have someone else find it. Here it is, and it is real.
In the Ahrefs breakdown, AI agents and agentic infrastructure sent 2,302 requests to llms.txt files, 10.5% of the total. AI retrieval bots sent 233, at 1.1%. The agentic use case is roughly ten times better evidenced than the search use case in the same dataset, from the same month, measured the same way.
That matches where the file came from. It was proposed for inference time assistance, and the companies serving it properly are documentation hosts: Stripe, Vercel, Cloudflare and Anthropic's docs domain all returned a real file in my audit, and all four of them are trying to make a coding agent's job easier, not trying to rank.
Google's Mueller drew the same line when Lily Ray asked him about serving markdown copies of developer docs.
OF COURSE they can read HTML just fine, so this is imo more of a temporary crutch, perhaps to save some tokens.
| As a search visibility lever | As a developer docs input | |
|---|---|---|
| Evidence in logs | 1.1% of requests, 233 in a month across 137K domains | 10.5% of requests, 2,302 in the same dataset |
| Engine commitment | None. Google states it does not use it | None formal, but agent tooling demonstrably fetches it |
| Mechanism | No referrer trail exists in any published log study | Token efficiency at inference, the original stated purpose |
| Honest verdict | Decoration | Reasonable if you ship a docs site |
So: if you run developer documentation, publish one. It costs an hour and there is measurable agent traffic. If you run a plumbing company, a law firm or an ecommerce store, the file will be read by your competitor's audit tool and nothing else. Ahrefs' Ryan Law put the structural problem better than I can.
I could also propose a standard (let's call it please-send-me-traffic-robot-overlords.txt), but unless the major LLM providers agree to use it, it's pretty meaningless
What I would do with the hour instead
The llms.txt hour is not the expensive part. The expensive part is that it occupies the slot where the mechanical work should go, and the mechanical work is unglamorous enough that almost nobody sells it.
- Confirm the retrieval crawlers can reach you. OAI-SearchBot, not GPTBot, is the gate on ChatGPT search answers. Blanket AI bot blocks and WAF challenge pages remove you from answer sets silently. This is the highest certainty lever available and it is the one most sellers skip. Full method in the AI crawler access audit.
- Check snippet eligibility, not just indexation. A stray nosnippet or max-snippet directive zeroes out Google AI feature eligibility on a page that looks perfectly healthy in Search Console.
- Read your own logs. Every claim in this post came from server logs. Yours will tell you which agents actually reach you, which is the only first party AI data you own. Start with AI crawler log analysis.
- Go get mentioned somewhere else. Citation weight overwhelmingly sits off your domain, which is why brand mentions beat on site markup work. The same logic that kills llms.txt also kills the schema for AI citations pitch.
- Measure before you change anything. Sampled, noisy, manual first. No first party share of voice exists inside any engine, whatever the tracking tools imply.
That is the ACCESS and MEASURE half of the Cited Method, and it is deliberately boring. Boring is the point. The novelty files get bought because they are legible to a client in a way that a robots.txt correction is not, which is a sales problem, not an engineering one. If you scope AI work for clients, the honest version of that conversation is covered in GEO retainer scope of work.
Check it yourself in five minutes
Do not take my word for any of this. The whole point of a log argument is that you have logs too.
- Request your own three files and compare: curl -s -o /dev/null -w "%{http_code} %{size_download} %{content_type}\n" https://yourdomain.com/llms.txt and the same for /robots.txt and /sitemap.xml
- Confirm the content type is text/plain and the body has real markdown in it, because HTTP 200 alone is not proof the file exists
- Grep your access logs for llms.txt over the last 30 days and count the hits
- Split those hits by user agent, then separate audit tools and tech profilers from named AI agents
- Compare that count against hits to /robots.txt over the same window, which is your control for whether crawlers were present at all
- If the llms.txt count is zero and the robots.txt count is not, you have reproduced every study cited on this page
My prediction, stated so it can be wrong: your llms.txt hit count for the last 30 days is under ten, and more than half of those hits are audit tools. If yours is different, publish it. This argument has been starved of data for two years and one more real dataset is worth more than another opinion piece, including this one.
The caveat I will keep repeating: every study here measures fetches and correlations, not causation, and none of them can see inside a model's retrieval stack. Nobody can prove llms.txt does nothing. What five datasets can show is that nothing is asking for it. On current evidence that is enough to stop selling it.
Frequently asked questions
Does llms.txt work for AI search visibility?
No. Three independent server log studies found near zero AI retrieval bot traffic to llms.txt files, and a roughly 300,000 domain analysis by SE Ranking found no correlation between having the file and being cited by AI models. No major engine has committed to reading it.
Do AI crawlers read llms.txt?
Barely. Ahrefs measured AI retrieval bots at 1.1% of all requests to llms.txt files across 137,210 domains in May 2026. Flavio Longato's Adobe Experience Manager log study independently reported the same 1.1% figure for verifiable LLM agents over a 30 day window.
Has Google said anything official about llms.txt?
Yes. Google's documentation changelog on 15 June 2026 states the files are not needed for Google Search and will not negatively or positively impact visibility or rankings, but that maintaining one for other systems is fine. That is permission, not endorsement.
Why do Anthropic, Stripe and Vercel publish an llms.txt file?
Because they run developer documentation, and Mintlify shipped llms.txt as a default across the docs sites it hosts. In my 28 domain audit, all four developer documentation hosts served a real file while OpenAI, Perplexity, Google and anthropic.com returned 404.
How many websites actually have an llms.txt file?
The published figures disagree by 13x. Ahrefs reported 28% inside its own analytics panel, SE Ranking found 10.13%, Longato measured 2.9% across 4,685 hosts, and HTTP Archive found 2.13% of desktop sites. The lower general web numbers are the more representative ones.
Is llms.txt worth adding anyway, since it is free?
If you ship developer documentation, yes, because agentic tooling demonstrably fetches it at roughly ten times the rate retrieval bots do. For a local business, ecommerce or professional services site, the file will mostly be read by audit tools and adds nothing.
Can llms.txt hurt my rankings?
No. Google states explicitly that the file will not negatively or positively impact visibility or rankings. The cost is opportunity cost, not risk. The hour spent generating and maintaining it is an hour not spent on crawler access, snippet eligibility or off site mentions.
How do I check whether anything is reading my llms.txt?
Grep 30 days of access logs for llms.txt, count the hits, then split them by user agent and separate audit tools from named AI agents. Compare that count against hits to robots.txt over the same window as a control that crawlers reached you at all.
Does llms-full.txt perform any better than llms.txt?
There is no published log study isolating llms-full.txt at scale that I could verify. Claims that models fetch it more often trace back to vendor commentary without disclosed fetch counts, so treat that as unmeasured rather than proven either way.
What should I do instead of installing llms.txt?
Confirm retrieval crawlers such as OAI-SearchBot are not blocked, verify pages are indexed and snippet eligible, read your own crawler logs, and earn brand mentions on third party properties. Those are the levers with documented mechanisms behind them.
Sources
- llmstxt.org (Jeremy Howard). The /llms.txt file proposal (2024-09)
- Ahrefs (Louise Linehan). We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read (2026-06)
- Flavio Longato. Do LLMs Read llms.txt? The Data from Adobe AEM Says Almost Never (2026-06)
- Otterly.ai. The llms.txt Experiment: 90 days of AI bot server logs (2026)
- SEO Depths (Simone De Palma). What Server Logs Tell About LLMS.TXT (2026)
- SE Ranking (Yulia Deda). LLMs.txt: Why Brands Rely On It and Why It Doesn't Work (2026)
- Google. Search Central documentation updates: Clarifying guidance on llms.txt files (2026-06)
- Google Search Central. AI Features and Your Website (2025-12)
- Google Search Central. Optimizing your website for generative AI features on Google Search (2026-07)
- HTTP Archive. Web Almanac 2025, SEO chapter (2026-01)
- Ahrefs (Ryan Law). What Is llms.txt, and Should You Care About It? (2026-06)
- PPC Land (Luis Rijo). llms.txt adoption stalls as major AI platforms ignore proposed standard (2025-07)
- Search Engine Journal (Matt G. Southern). 97% Of llms.txt Files Got No Requests, Ahrefs Data Shows (2026-06)
- Search Engine Journal. Mueller Explains Why Google Uses Markdown On Dev Docs (2026-05)
- Cameron Rye. The llms.txt Standard: Why Nobody Uses It (2025-09)
- Ahrefs (Louise Linehan). AI search trends 2026 (2026-07)
Want me to run this on your site and show you the before and after?
One call, no pitch deck. We look at what is actually blocking you and tell you the truth about whether we can help.
Book a free consultation →