Updated
What llms.txt is
llms.txt is a Markdown file served at the root of your domain, at /llms.txt. It gives a language model a short description of the site and a curated list of links to the pages that matter most, ideally to clean text or Markdown versions of them.
The idea is practical: a model has a limited context window, and a typical web page is mostly navigation, scripts and layout. A hand-picked index gets it to the useful content faster.
The format
The proposal keeps the structure simple so both people and programs can read it:
- An H1 with the name of the site or project. This is the only required part.
- A blockquote with a one or two sentence summary.
- Optional paragraphs with further context.
- H2 sections, each containing a Markdown list of links in the form [title](url): short note.
- An optional section titled Optional for links that can be skipped when context is tight.
llms.txt is not robots.txt
The two files do different jobs. robots.txt controls access: it tells crawlers, including AI crawlers such as GPTBot, ClaudeBot and PerplexityBot, which paths they may fetch. llms.txt grants and blocks nothing. It is a reading guide, not a permission system.
If you want to keep AI crawlers out, or let them in, that is still done in robots.txt.
Does it actually help?
llms.txt is a community proposal, not a standard, and no major AI provider has committed to using it for ranking or citation. Treat it as a low-cost, possibly useful signal rather than a lever with a guaranteed effect.
What clearly does matter for being cited by AI assistants is the same groundwork that matters for search: pages that are crawlable, content that is present in the HTML rather than rendered late by JavaScript, accurate structured data, and AI crawlers that are not blocked by accident.
Check your own site
The GEO checker looks at all three foundations at once: whether an llms.txt file exists, whether robots.txt lets AI crawlers in, and whether the page carries structured data.