Guide · Content & technical

What Is llms.txt and Does Your Site Need One?

llms.txt is a proposed Markdown file that gives AI tools a curated map of your site. Here's the format, what it can and can't do for AI visibility, and when it's worth publishing.

By InTheAnswer Editorial · Updated · 8 min read

llms.txt is a proposed Markdown file, placed at the root of a site as /llms.txt, that gives AI tools a short summary of the site and a curated list of links to its most useful pages. Most sites don't need one to be cited by answer engines: none of the major engines has documented using it to decide what to retrieve or cite, and Google says Search ignores it.

That doesn't make it pointless. It's cheap to publish, it gives AI agents and coding assistants a clean entry point to documentation, and writing one forces you to describe your business in a few precise sentences, which helps everywhere else too. It just isn't a shortcut to AI citations.

This guide covers where llms.txt came from, what the format looks like, what it is and isn't, an honest cost-benefit view, and how InTheAnswer publishes its own.

Where llms.txt came from

The proposal was published at llmstxt.org by Jeremy Howard of Answer.AI on September 3, 2024. Its starting point is a practical problem: web pages are built for people, so their content is wrapped in navigation, ads and scripts that are hard to convert into clean text, and context windows are still too small to hold most websites whole. A short Markdown file at a predictable address gives a language model the essentials without the clutter.

The idea borrows its location from /robots.txt and /sitemap.xml, but not their purpose. The proposal itself draws the line: robots.txt tells automated tools what access is acceptable, while llms.txt is meant to be read on demand when someone is using a model to understand a site. Its author expected it to be used mainly at that moment of use, called inference, rather than for training.

A second version followed in August 2026, which the author says reflects two years of adoption. According to the proposal's site, thousands of sites now publish an llms.txt file and documentation platforms generate one automatically.

What the format looks like

The format is ordinary Markdown in a fixed order. Only the first element is required.

  1. An H1 with the site or project name. This is the only mandatory part.
  2. A blockquote with a short summary of what the site is and the key facts needed to understand the rest.
  3. Optional free text, such as paragraphs or lists with important notes.
  4. H2 sections of links, each a Markdown list of links with an optional note after a colon.
  5. An "Optional" section, by convention, for secondary links an agent can skip when it needs to keep things short.

Here's what that looks like for a fictional analytics company:

# Example Analytics

> Example Analytics is a product analytics tool for B2B software teams. This file lists the pages that best explain what it does, what it costs and how to set it up.

Key facts:
- Plans are billed monthly per workspace.
- Customers can choose EU or US data hosting.

## Product
- [Features](https://www.example.com/features.md): event tracking, funnels and retention reports
- [Pricing](https://www.example.com/pricing.md): plans, limits and billing terms

## Docs
- [Quick start](https://www.example.com/docs/quick-start.md): install the tracking snippet and send a first event
- [API reference](https://www.example.com/docs/api.md): endpoints, authentication and rate limits

## Optional
- [Changelog](https://www.example.com/changelog.md): release notes by month

The proposal also suggests offering clean Markdown versions of pages at the same URL with .md added, which is why the links above end that way. Version two adds standard ways for a page to point to its Markdown version and to the llms.txt that covers it, and lets a file in a subfolder such as /docs/llms.txt cover just that section. None of this is required; plain links to normal HTML pages still work.

What llms.txt is, and what it isn't

It's easy to load more onto this file than it can carry.

  • It is a curated index. It's the page you'd hand a new colleague: what we are, and where the important information lives.
  • It isn't an access control. It doesn't allow or block anything. A crawler blocked in robots.txt stays blocked whatever llms.txt says, and a crawler allowed in robots.txt doesn't need llms.txt to read your pages. How to configure robots.txt for AI crawlers covers the file that does control access.
  • It isn't a sitemap. A sitemap lists every URL for search engines; llms.txt is short, selective and can link to other sites.
  • It isn't a documented ranking or citation signal. As of this writing, none of the major answer engines (ChatGPT, Perplexity, Claude, Gemini, Google AI Overviews or Copilot) has publicly documented using llms.txt to decide which pages to retrieve or cite.
  • It isn't something Google uses. Google's guide to optimizing for generative AI search says you don't need machine-readable files, AI text files, markup or Markdown to appear in Search, including its AI features. It adds that publishing llms.txt for other services is fine and will neither help nor harm your Google visibility.
Key takeaway: llms.txt is a convenience for AI tools that choose to read it, not a lever on how answer engines rank or cite. Crawler access, indexation and coverage on trusted sources still decide whether you're cited.

Who actually reads it

The proposal expects agents to open llms.txt, find what they need and follow the links. The clearest real-world fit is documentation: a developer working with an AI coding assistant can point it at a product's llms.txt and get the relevant docs without the assistant crawling an entire site. The fact that documentation platforms generate the file automatically suggests that's where most of the adoption sits.

Answer engines retrieving pages for a consumer question are a different case. As covered in how AI answer engines choose sources, they search an index and read the pages it returns. Nothing in their public documentation says that process consults llms.txt. A file nobody has said they use may still be read occasionally, for example when a user pastes your URL into a chat, but that's a guess, not a plan.

The honest cost-benefit

The costs are small, but they aren't zero:

  • Writing time. A useful file takes some thought about what matters most. An automatic dump of every URL defeats the point.
  • Maintenance. A file listing old prices or retired products is worse than none, because it states outdated facts in a format built for machines to trust.
  • Exposure. Everything in it is public. Don't list staging URLs or pages you'd rather not have read.

The benefits are modest but real:

  • A clean entry point for agents and assistants working with your docs or product information.
  • A forcing function. Writing a two-sentence summary of what you do, for whom, is the same exercise that makes your About page, schema and third-party descriptions consistent.
  • Low risk. By Google's account it neither helps nor harms Search, and no engine has said it penalizes one.

Our view by site type:

  • Developer tools, APIs and software with extensive docs: worth publishing, since this is what the format was designed for.
  • SaaS and e-commerce with complex products or pricing: reasonable if you'll keep it current; low priority if you won't.
  • Local businesses, brochure sites and small publishers: optional. Time spent on crawler access, indexation and accurate listings will do more.

Whatever you decide, don't let it displace the work that does affect citations: fixing crawler access, writing pages engines can quote, and being described on publications they already trust.

How to publish one well

If you do publish one, a few habits keep it useful:

  1. Write the H1 and blockquote first, using the same wording you use on your About page and in your structured data.
  2. Link to the pages you'd point a new customer or colleague to, not to everything.
  3. Add a short note after each link saying what the page answers.
  4. Move secondary material into an "Optional" section.
  5. Serve it as UTF-8 plain text at /llms.txt, make sure it returns a 200 status, and check that your firewall doesn't block AI user agents from reading it.
  6. Generate it from the same source as your site where you can, so prices and product names can't drift.

How InTheAnswer publishes /llms.txt and /llms-full.txt

InTheAnswer publishes a /llms.txt file that follows the proposal's structure: an H1, a blockquote describing the marketplace, key facts about the catalog, the AEO scoring method, and H2 sections linking to core pages, every guide, each industry page and the top 50 featured placements. Its Optional section links to the sitemap and to /llms-full.txt, a longer version that adds how prices work, how to describe and cite InTheAnswer accurately, and every featured placement with its metrics. An llms-full.txt file is a companion convention some sites use; it isn't part of the llmstxt.org proposal itself.

Both files are generated from the same catalog data as the website, so their figures update whenever the catalog does; it currently holds 10,397 placements. Our robots.txt explicitly allows the major AI crawlers to read both. A human-readable version of the same guidance lives on the AI instructions page.

We publish these files because they're cheap to maintain and keep our facts consistent wherever a model reads them, not because any engine has said it ranks sites by them. If you want to check what actually matters for your own pages, the free AEO audit reports AI crawler access, structure and schema for any URL.

Frequently asked questions

Does llms.txt help me get cited by ChatGPT or Perplexity?+

There's no documented evidence that it does. Neither OpenAI nor Perplexity has said its search systems use llms.txt, and both document crawlers that read ordinary web pages. Crawler access, indexation and presence on sources those engines retrieve are what their documentation points to.

Should llms.txt replace robots.txt or my sitemap?+

No. Robots.txt controls which crawlers may access which paths, and a sitemap helps search engines find every URL. llms.txt does neither; it's a curated summary that sits alongside both.

What's the difference between llms.txt and llms-full.txt?+

llms.txt is the short, curated file the proposal defines. llms-full.txt is a companion convention some sites use, not part of the proposal, that bundles more of the actual content into one file so a tool can read it without following links. Publish it only if you can keep both accurate.

Can I use llms.txt to stop AI companies training on my content?+

No. llms.txt has no opt-out mechanism. Training opt-outs are handled in robots.txt with tokens such as GPTBot, ClaudeBot, Google-Extended and Applebot-Extended, each of which its company documents.

From the catalog

Top Tech / SaaS placements by AEO score

All Tech / SaaS links →
EliteAEO 96.0 Verified
Tech / SaaS
DR
97
Ahrefs
DA
96
Moz
Traffic
46M
Ahrefs
United States 24%·Nofollow
$489/ placement
EliteAEO 89.8 Verified
Tech / SaaS
DR
93
Ahrefs
DA
—
Moz
Traffic
4.4M
Ahrefs
United States 46%
$2,100/ placement
EliteAEO 88.0 Verified
Business / FinanceTech / SaaS
DR
92
Ahrefs
DA
94
Moz
Traffic
5.7M
Ahrefs
India 93%·Nofollow
$1,250/ placement
Keep reading