bettersorted Logo
Laptop with website code – symbolic image for technical optimization
AI Consultancy

llms.txt, Schema.org & Co.: How your website is cited by AI search systems

AuthorMuhamed Alahmed
Published on
Reading time5 min
In short

For ChatGPT, Perplexity or Google AI Overviews to cite your website, three things matter most: crawlable, fast pages, clearly structured content with direct answers, and clean structured data (Schema.org). An llms.txt can be added as a supplement, but it is not an official standard – according to Google, it is not used.

The most important points at a glance

  • AI search systems cite pages that they can access, understand, and classify as trustworthy.
  • Three levels: crawlability, answer-first content, Schema.org data.
  • llms.txt is not a standard; according to Google, it is not used.
  • Structured data must match the visible content – otherwise there will be errors.

AI search systems such as ChatGPT, Perplexity or Google AI Overviews cite websites that they can technically access, understand in terms of content, and classify as trustworthy. Three levels matter for this: crawlability and speed, clearly structured content with direct answers, and structured data according to Schema.org. A llms.txt can be a useful addition, but it is not an official standard – and Google has stated that it does not use it.

This post is the technical companion to GEO instead of SEO and the GEO practical guide. What GEO fundamentally means is explained in the glossary entry Generative Engine Optimization.

Level 1: Accessibility for AI crawlers

Before an AI system can cite your content, it must be allowed and able to read it. Check three points:

  • robots.txt: Are crawlers such as GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot or Google-Extended blocked? Many websites block them unintentionally through blanket rules.
  • Server-side rendering: Content that is only generated in the browser via JavaScript is not seen by many crawlers. Important text should be included directly in the delivered HTML.
  • Speed and accessibility: Slow or faulty pages are crawled less often and less completely.

An example from our own website: expandable FAQ answers used to be displayed only via JavaScript – as a result, crawlers saw only a fraction of the answers. Since switching to native HTML elements, all answers are directly present in the source code.

Person researching on a laptop

Level 2: Content that can be quoted

AI systems prefer to adopt passages that answer a question directly and in a self-contained way. This leads to simple writing rules:

  • Answer first:The core question is answered in the first two or three sentences, followed by a more detailed explanation.
  • Questions as subheadings:The way people would ask them to a chatbot.
  • Specific facts with source:Figures, deadlines, conditions – substantiated and with a date of reference.
  • Lists and tables:Compare and structure steps instead of hiding them in running text.
  • Clear terms:Explain technical terms once, then use them consistently.
  • Recognizable authorship:Who is writing, and with what experience? This is a signal of trust.

Level 3: Structured data according to Schema.org

Structured data describe content in a machine-readable way in JSON-LD format. They help search engines understand who offers something, what a page is and which questions it answers. Important types for company websites:

  • Organization / LocalBusiness: Name, address, contact, profiles in other networks (sameAs).
  • Person: Authors with qualifications.
  • Service:Services with provider and area of use.
  • Article / BlogPosting:Posts with author, date, and update.
  • FAQPage:Questions and answers.
  • BreadcrumbList:the classification of a page within the website structure.

Important: Structured data must match the visible content and be complete. Product markup without a price, for example, generates errors in Google Search Console – if you do not publish prices, it is better to describe your offerings as a service.

And the llms.txt?

The llms.txt is a text file proposed in 2024 in the root directory of a website that explains to AI systems in compact form what the website is about and which pages are the most important – more in the glossary under llms.txt. It is not an official standard. Representatives from Google have publicly stated that Google Search does not use it; whether and how other AI providers evaluate it is not reliably documented.

Our assessment: an llms.txt does no harm and can be created quickly, but it should not be the first measure. Crawlability, good content, and correct structured data have a proven effect – llms.txt is a small addition.

Checklist in seven points

  1. Check robots.txt for blocked AI crawlers.
  2. Important content in the delivered HTML instead of only via JavaScript.
  3. One clear H1 per page and questions as subheadings.
  4. Put the answer first in the opening sentences.
  5. Check Organization, Service, Article, and FAQPage schema and fix errors in Search Console.
  6. Make authors visible, and include the date and last updated information.
  7. Optional: create llms.txt with the most important pages.

How well your own website is set up for AI search systems is shown by the free SEO/GEO check in just a few minutes.

How an llms.txt is structured

The proposal calls for a simple Markdown file at /llms.txt:

  • a heading with the name of the website or company
  • a brief summary in one paragraph
  • sections with lists of the most important pages, each with a link and one sentence of description
  • an optional more detailed version (llms-full.txt) with more content

Anyone creating an llms.txt should keep it up to date – outdated links and descriptions are worse than no file at all. bettersorted.de provides such a file, but considers it an addition, not the core of optimization.

E-E-A-T: why trust signals matter

Google describes with E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness) the signals used to assess content quality. Similar standards apply to AI search: systems prefer to cite sources they trust. Specifically, the following help:

  • Recognizable authors with profiles and qualifications
  • Practical relevance: own experience, examples, references
  • References: Sources and date of publication for facts
  • Legal notice, contact, privacy policy complete and easy to find
  • Mentions on other websites: directories, specialist portals, press, networks

How to measure whether AI systems cite your site

  1. Ask the typical questions of your target audience in ChatGPT, Perplexity, and Google, and check which sources are cited.
  2. In web analytics, monitor visits from AI services (referral domains).
  3. In Google Search Console, track impressions for question-based search queries.
  4. Repeat the check at regular intervals, because answers change.

Common technical errors

  • AI crawlers accidentally blocked via robots.txt or firewall
  • Important content only in images or PDFs
  • Answers in expandable elements that are only loaded via JavaScript
  • Incorrect or incomplete structured data
  • Duplicate pages without a canonical tag

Frequently asked questions

What is an llms.txt?

A text file in the root directory of a website that explains to AI systems in a compact form what the website is about and which pages are important. It was proposed in 2024 and is not an official standard.

Does Google use llms.txt?

No. Google representatives have stated that Google Search does not use llms.txt. For Google, crawlability, content, and structured data matter.

Which structured data is most important for AI search?

Organization or LocalBusiness for the company, Person for authors, Service for services, Article for posts, FAQPage for questions, and BreadcrumbList for the site structure.

Should I allow AI crawlers in the robots.txt?

If you want your content to appear in AI responses, yes. Anyone who blocks GPTBot, ClaudeBot, or PerplexityBot cannot be cited by these systems. The decision is a trade-off between visibility and control over your own content.

As of October 2026.

Portrait of Muhamed Alahmed, founder of bettersorted
About the author

Muhamed Alahmed

With over 10 years’ experience in IT, I develop solutions that not only work from a technical perspective, but also create real added value and open up new possibilities.

More about bettersorted →

More articles

Newsletter

Stay up to date on AI topics

Short updates on AI automation, funding programs and new posts — no spam, unsubscribe anytime.

We use Brevo as our marketing platform. By submitting, you agree that your data will be transferred according to Brevo's privacy policy .

CallEmailContact form