You Don't Need an llms.txt File
AI Marketing

You Don’t Need an llms.txt File (And the Data Explains Why)

By, Carlos Rios
  • 15 Sep, 2026
  • 86 Views
  • 0 Comment

Ninety-seven percent of llms.txt files on the web were never fetched by anything at all in May 2026. Not a crawler. Not a bot. Not a human. That number comes from Ahrefs’ server-log analysis of 137,210 domains, published 15 June 2026, and it is the closest thing this industry has to a verdict on a file that agencies have been selling as essential for the better part of two years.

If you are holding a quote with “llms.txt implementation” as a line item, this is your second opinion.

We are not neutral here. We sell generative search work. Recommending against a billable deliverable is not a great commercial instinct, but the evidence on this one is not close, and we would rather you spend the money on something that moves.

You do not need an llms.txt file to appear in AI search. Google’s own documentation lists it as unnecessary, and server logs from 137,000 domains show 97% of published files are never requested by anything. The file has one genuine use case: developer documentation consumed by AI coding agents. For everyone else, it is decoration.

What is an llms.txt file, and who actually proposed it?

An llms.txt file is a single markdown file placed at a website’s root that summarises what the site is and links to its most important pages. The idea is that a language model could read it and orient itself without crawling everything.

It was proposed in September 2024 by Jeremy Howard, co-founder of Answer.AI and fast.ai, and published at llmstxt.org. Two details about its origin get lost in most coverage, and both matter.

First, it was designed for developer documentation. Howard’s use case was AI coding tools needing to parse an API reference quickly, not marketing sites trying to get quoted in ChatGPT. The AI visibility framing was attached later, by the SEO industry, on speculation.

Second, despite the filename, it is not a robots.txt-style directive. It controls nothing and blocks nothing. It is a suggestion in a file that no major platform has agreed to read. There is no IETF RFC behind it and no W3C working group. It remains a community proposal.

Neither of those facts makes it a bad idea. They just make it a proposal, and proposals need adoption to become standards.

Does Google use llms.txt?

No, and Google has now said so in writing. On 15 May 2026, Google Search Central published Optimizing your website for generative AI features on Google Search. It contains a section titled “Mythbusting generative AI search,” and llms.txt is named in it. Google’s position is that its crawler may discover such files, but treats them like any other text file. No special handling. No preferred indexing pathway.

This was not a reversal. Gary Illyes told Search Central Live attendees in July 2025 that Google does not support llms.txt and has no plans to. John Mueller had already compared the concept to the keywords meta tag, the 1990s signal search engines abandoned because site owners controlled it and therefore gamed it.

There is one wrinkle worth knowing about, because it is the source of most of the confusion you will encounter.

Days after that guide went live, Chrome shipped an llms.txt check inside Lighthouse’s experimental agentic browsing audits. Two Google teams, two apparent positions, one very loud week on LinkedIn. When Lily Ray asked Mueller about the contradiction directly, his answer was that llms.txt is “not done for search.” He described it as a temporary crutch for AI coding tools parsing developer documentation.

That is the whole contradiction, resolved: the Chrome team is thinking about agents reading documentation. The Search team is thinking about ranking. Different problems, different answers, and only one of them is about whether you show up in an AI Overview.

Do AI crawlers actually read llms.txt files?

Barely, and the specific bots you care about read them least. This is where the Ahrefs study earns its place, because until it was published, everyone in this debate was arguing from single-site logs and vibes.

The methodology: 137,210 domains with measurable traffic in May 2026, checked for a valid llms.txt returning HTTP 200, then cross-referenced against every request hitting /llms.txt paths, classified by user agent.

Twenty-eight percent of those domains published a file. Ahrefs flags that as an upper bound, since their analytics customers skew technical and SEO-aware. Of those roughly 38,000 files, 97% received zero requests in May. The remaining 3%, about 1,100 domains, absorbed every request measured.

Inside that small pool, the breakdown is the part worth sitting with. AI retrieval bots, meaning OAI-SearchBot, PerplexityBot and Claude’s search crawler, accounted for 1.1% of requests. Those are the bots that fetch pages to answer live user queries. They are the entire reason anyone publishes this file for visibility purposes.

Slackbot fetched llms.txt files more often than PerplexityBot did. Slackbot is a link-preview unfurler. It was almost certainly SEOs pasting llms.txt URLs to each other in chat.

The largest single category of requests, at 21.7%, was SEO audit tools checking whether the file existed.

Bot request bar chart

So who is reading them?

Mostly the industry that invented the debate, plus one group nobody was optimising for.

Twelve percent of all requests came from tools studying the standard rather than consuming it: GEO and AEO scoring platforms at 5.8%, dedicated llms.txt scanners and validators at 3.6%, and research crawlers at 2.7%. An ecosystem of auditors formed around a file before anyone established that the intended readers had arrived.

The genuine readership is AI agents and the infrastructure built to serve them, at 10.5% of requests. Anthropic’s Claude Code out-fetched every AI retrieval bot, every AI assistant and every training crawler in the dataset. Which is precisely the use case Howard designed it for, and precisely what Mueller described.

One finding deserves more attention than it has received. The largest single research crawler in the data identified itself as prompt-injection-survey/1.0. Someone is systematically probing llms.txt files as an attack surface, on the reasonable theory that agents are built to trust them. If you publish one, treat it like code: version control it, restrict edit access, keep it to plain links and descriptions with nothing instruction-shaped, and review whatever your CMS auto-generates on your behalf. A stale or compromised file misleads every agent that reads it.

And then the finding that ends the “but what if I am missing out” argument: across every 404 response to /llms.txt, the AI bot share was zero. No AI system probes for a file you have not published. The people checking for absent llms.txt files were humans typing URLs into browsers, presumably competitors. Nothing is knocking at a door you have not built.

Why did this become an agency upsell?

Because it is easy to sell and impossible to disprove on a short timeline.

An llms.txt file takes under an hour to produce and produces a deliverable the client can look at. AI search visibility takes months and depends on things clients find harder to buy: genuine expertise on the page, technical crawlability, and an entity Google and the model providers can actually resolve. One of those is a tidy invoice line. The other is a programme.

Why did this become an agency upsell?

Website builders made it worse by turning it into a default. Wix generates these files. Framer and Lovable scan for them. Adoption arrived as a platform setting before it was ever a webmaster decision, and platform defaults look a lot like consensus from the outside.

We would put this in the same category as the other tactics Google named in that same mythbusting section: content chunking, AI-specific rewrites, bought mentions, and over-application of structured data. None of them are frauds exactly. They are all plausible-sounding work that produces an artifact, which is a different thing from producing a result. This is the distinction the Marketing Ownership Framework is built around: you should be able to look at any line on an invoice and say what it changed.

When is llms.txt actually worth publishing?

Three situations, and only three.

You run developer documentation. If your customers point coding assistants at your API reference, a well-formed llms.txt and llms-full.txt gives those tools clean, current documentation instead of a guess. This is the original use case and the only one the data supports today.

Agents transact on your site. If AI agents act on your behalf or on your customers’ behalf inside your product, helping them orient has direct value that has nothing to do with search rankings.

Your CMS does it for free. If Wix or a plugin generates the file without your involvement, there is no reason to remove it. It is neutral for Google. Just link to it from your HTML so agents can find it, because agents fetch llms.txt when directed, not speculatively, and then govern it like code.

What none of those situations is: a B2B services site publishing a file in the hope that Perplexity notices.

What to do instead

The uncomfortable answer is that there is no file-shaped shortcut, and the things that work are the things that were already working.

Google’s own guidance, in the same document that dismissed llms.txt, points at non-commodity content: material with a first-hand point of view that could not have been assembled from other pages. AI systems construct answers from sources. Being one of those sources means having something the model cannot synthesise from the other eleven results.

Then the boring layer underneath it. Confirm your priority pages are crawlable and indexed. Confirm your robots.txt is not blocking GPTBot, ClaudeBot, PerplexityBot or Google-Extended, because if those agents cannot reach the page, nothing else on this list matters. Get Organization and sameAs schema live sitewide so the model providers can resolve who you are as an entity rather than as a string.

We have written the longer version of this argument in how to get your brand cited by AI and covered the mechanics of ranking in AI search engines separately. The short version: the work that earns AI citations is the work that earns links and rankings, done to a higher standard, on a domain with a resolvable identity.

None of it fits in a text file.

Frequently asked questions

Does Google read llms.txt?

Google’s crawler may discover an llms.txt file, but treats it as an ordinary text file with no special handling. Google Search Central’s May 2026 generative AI guide names llms.txt in its mythbusting section, and Gary Illyes confirmed at Search Central Live in July 2025 that Google does not support it and has no plans to.

Does ChatGPT use llms.txt?

OpenAI has never committed to reading llms.txt at retrieval time, and its crawler documentation points to robots.txt instead. In Ahrefs’ May 2026 log study, GPTBot was the single most active fetcher among AI training crawlers at 4.51% of requests, but OAI-SearchBot, the bot that answers live queries, barely registered.

Will an llms.txt file hurt my SEO?

No. Google has stated the file neither helps nor harms rankings. The realistic cost is the time to create and maintain it, plus one security consideration: researchers have begun probing llms.txt files as a prompt-injection surface, since agents are designed to trust them. A stale or compromised file misleads every agent that reads it.

What should I do with an llms.txt file my website builder created?

Leave it, but govern it. Link to it from your HTML so agents can discover it, since agents fetch the file when directed rather than speculatively. Keep the contents to plain links and descriptions with no instruction-shaped text, restrict who can edit it, and review anything the platform regenerates automatically.

Is llms.txt the same as robots.txt?

No. Robots.txt is a formally standardised protocol (RFC 9309) that crawlers use to determine which parts of a site they may access. Llms.txt is a community proposal with no standards body behind it, and it controls nothing. The shared naming convention has caused most of the confusion around what the file can actually do.

Who should publish an llms.txt file?

Developer documentation sites whose users point AI coding assistants at their API references, and products where AI agents act on behalf of users. For those cases, the Ahrefs data shows real fetch activity from agentic tools. For marketing and services websites pursuing AI search visibility, the evidence does not support it.


If your agency has quoted you for AI search work, ask which line items Google has named as unnecessary. That one question tends to clarify the rest of the proposal quickly.

Our generative engine optimisation work starts with the entity and technical layer, then the content that gives a model a reason to cite you. No files that nothing reads.

Want a second opinion on your AI search proposal?

We will tell you which parts of it Google has already named as unnecessary, and what the same budget buys instead.

Book a 20-minute review

Carlos Rios

Author

Carlos Rios

Carlos Rios is the Founder of Tabula. Before starting the agency, he spent his career inside companies — reporting to leaders who had no patience for vanity metrics and wanted straight answers about what was working and why. He built Tabula around that same standard: a marketing system the client owns, powered by AI and led by experts who explain the plan in plain language. Carlos studied philosophy for six years at York University and holds a Master's in Marketing from the Schulich School of Business.