Generative search changes who reads your page first. Increasingly it is a retrieval system deciding whether your content can be cited, summarised or ignored. That system is not impressed by tone, and it cannot see your art direction.
What actually helps
- A clear entity definition: what the organisation is, in one sentence, in the same words everywhere.
- Answers placed immediately under the heading that asks the question — not four paragraphs later.
- Factual statements with a scope, rather than superlatives with none.
- Structured data that matches the visible page exactly.
- Headings that describe content instead of performing cleverness.
None of this is new. It is the same advice technical SEO has given for a decade, with the tolerance for vagueness removed. What has changed is the cost of ignoring it: a page that never states plainly what the company does can still rank, but it cannot be quoted.
The one genuinely new requirement
Consistency across surfaces now has retrieval value. If the site, the structured data, the profiles and the press all describe the organisation slightly differently, a retrieval system has several competing descriptions and no reason to prefer yours. Pick one sentence and use it everywhere, including in the schema.
What does not help
Text hidden from users but served to crawlers has always been a violation and still is. So is structured data that claims something the page does not show — a rating with no reviews, an FAQ block with no visible FAQ. Both are detectable, both are penalised, and both are the kind of shortcut that costs more to unwind than it ever returned.
A simple test: read your markup as plain text with the CSS removed. If a stranger cannot tell what the company does, what it sells and who it is for, no retrieval system will either.
What to mark up, and what to leave alone
Structured data is a description of the page, not an enhancement of it. Mark up the things that are genuinely on the page and genuinely typed: the organisation, the page itself, the breadcrumb trail, an article and its dates, a service and what it delivers, and an FAQ the reader can actually read. That set covers almost every business site.
Leave the rest alone. A schema type that is technically applicable but describes something the page does not show buys nothing and creates a mismatch that is expensive to unwind. The useful test is whether a person reading the page could verify the markup by looking at it.
How to tell whether it is working
Classic rank tracking will not show you this. What shows it is the shape of the traffic: fewer sessions for a given query set, arriving with more intent, because the summary answered the browsing questions and forwarded only the people who wanted the source. Treat a drop in informational traffic alongside stable enquiry volume as a signal, not a loss.
- Check whether your own definition sentence is the one being repeated back to you.
- Watch referrals from assistant and answer surfaces separately from classic search.
- Validate structured data on a schedule, not once at launch — it rots quietly when content changes.
- Keep a list of the questions the business should be the answer to, and re-read the pages that answer them.
Where llms.txt fits
A machine-readable summary at a predictable path is cheap to maintain and useful as a statement of intent, but it is not a ranking mechanism and no major system treats it as authoritative. Publish one if it reflects the real pages. Do not publish one that says something the site does not.