Free · no account · transparent crawl

Create an llms.txt that stays useful

Crawl a public website, select the pages that matter, edit the Markdown and validate the published file. Then keep it healthy with a free weekly 0–100 score and three concrete next steps.

The generator currently uses a German interface, but it can crawl and structure websites in other languages. The English product interface is being expanded.

Weekly scorecard

86/100
  • Add a traceable “last reviewed” date.
  • Remove two duplicate links.
  • Fix one sitemap URL returning HTTP 404.

A shortlist, not another sitemap dump

llms.txt is a proposed Markdown format for presenting a small, curated map of a website. A strong file names the authoritative product, documentation, pricing, policy and expertise pages. It does not contain crawler permissions, guarantee citations or replace source-page quality.

Curate

Prioritise the few pages that explain your organisation and offering best.

Validate

Check structure, links, response headers, duplicates, file size and review date.

Maintain

Receive a weekly scorecard before broken links and stale references accumulate.

Three files, three jobs

FilePurposeWhat it does not do
llms.txtCurates important public sources with context.No crawler access control and no ranking guarantee.
robots.txtDefines voluntary crawl rules by user-agent and path.No authentication and no content prioritisation.
sitemap.xmlLists canonical, indexable URLs for discovery.No guarantee of indexing, ranking or AI citation.

Transparent by design

The crawler is deliberately bounded. It reads public HTML, robots.txt and sitemaps, remains on the same registrable domain, blocks private network targets and does not execute page JavaScript. The output stays editable because curation is an editorial decision.

  1. 1Enter a public domain and run the bounded crawl.
  2. 2Review selected and excluded pages instead of accepting a blind URL dump.
  3. 3Edit descriptions, group sources and validate the Markdown.
  4. 4Publish /llms.txt with HTTP 200, ideally as text/plain.
  5. 5Activate weekly monitoring and fix the highest-impact issue first.

The practical USP

Publishing is easy. Staying accurate is valuable.

The free monitor checks your published file, sitemap, status codes, canonicals and a limited set of important pages every week. The email leads with one understandable score and the three most useful next actions. Promotional GEO or WebMCP information is only included after a separate, optional consent.

Get your weekly score

After double opt-in, you receive your current llms.txt score, the most useful improvements and noteworthy technical changes every week.

Frequently asked questions

Is llms.txt an official web standard?

No. It is an open proposal, not an adopted IETF, W3C or ISO standard. Support is still product-specific and may change.

Does llms.txt control AI crawlers?

No. Use robots.txt for cooperative crawler access rules and real authentication for private content. llms.txt is a curated content guide.

Will it make ChatGPT or Perplexity cite my website?

There is no guarantee. Clear source pages, verifiable claims, crawlability, stable canonicals, authorship and authority matter far more than one file.

How many pages should the file include?

Use a deliberate shortlist of the pages that best explain your company, products, documentation, policies and expertise. A complete URL inventory belongs in a sitemap.

What does the quality score measure?

It checks technical signals such as reachability, HTTP response, content type, headings, summary, Markdown links, duplicate or external URLs, file size and review date. It is not a ranking score.

How often should I update llms.txt?

Review it when products, pricing, documentation, policies or key URLs change. Weekly monitoring catches technical drift while editorial decisions remain yours.

Do you send my crawled content to an AI provider?

No. The current generator extracts and structures public HTML deterministically and does not send the URL or crawled content to an external language model.

Where is monitoring data stored?

The production Redis primary region is configured in AWS eu-central-1, Frankfurt. The privacy policy explains the providers, retention logic and possible international transfers.

Primary references: llmstxt.org, RFC 9309 and Google's sitemap documentation.

Start with the file. Continue with GEO and agent readiness.

Use GEO-Tool for broader technical visibility checks and WebMCP-Tool for the next layer: turning discoverable content into safe, declared agent actions.