Adding Schema Didn't Move AI Citations. Here's What Does.

Somewhere in the last two years a checklist appeared. Add FAQ schema markup. Publish an llms.txt file, a plain-text summary written for AI crawlers. Cut your pages into smaller pieces so a model can digest them. Whole product categories now sell against that list, and most content teams have worked through part of it.
Enough of it has now been measured to say something useful. The largest study tracked 1,885 pages that added schema markup against 4,000 matched control pages, and found no gain on any engine it looked at. A separate test suggests why: when five AI systems were asked to fetch a page and report a price hidden in its markup, not one of them could.
What follows is what the evidence supports, where it runs out, and how to check the next statistic someone quotes at you.
Key Takeaways
- Ahrefs tracked 1,885 pages that added JSON-LD schema against 4,000 matched control pages: Google AI Overviews fell 4.6%, AI Mode rose 2.4%, ChatGPT rose 2.2%. Only the small decline was large enough to be real.
- On a test page where prices were hidden in markup, none of ChatGPT, Claude, Perplexity, Gemini or Google AI Mode retrieved them. What each took from the visible page varied widely, and Claude took nothing at all.
- Optimising page text can backfire. Run through a full search pipeline over 170,000 documents, body-only rewrites cut a page's chance of being retrieved by about 9% and its final citation by 6%, while improving how much a retrieved page gets quoted.
- Tested across six domains, only 3 of 54 method-and-domain pairings of popular optimisation tactics produced a real gain, and several pushed pages down the rankings.
- Across 75,000 brands, branded web mentions track AI Overview visibility about three times more closely than backlinks do, 0.664 against 0.218 on a 0-to-1 scale.
- Do this instead: put explicit facts on your top pages. Prices, specifications, dates, named numbers. Across 252,000 trials and six models, a missing price was one of the two strongest negative signals.
Does Schema Markup Increase AI Citations?
On the largest measurement published, no. Across three engines it produced no gain at all, and on Google AI Overviews a small decline. A separate test suggests why: nothing hidden in markup gets read when an engine fetches your page.
Schema markup is JSON-LD: a block of code in your page source, invisible to readers, that labels parts of the page for machines. FAQPage markup tells a crawler "this is a question and this is its answer."
Ahrefs, an SEO tool company, tracked 1,885 pages that added JSON-LD schema between August 2025 and March 2026. Each was compared against control pages that changed nothing, 4,000 of them, matched on how often they were already being cited.
| Where | Change in citations, against controls |
|---|---|
| Google AI Overviews | -4.6% |
| Google AI Mode | +2.4% |
| ChatGPT | +2.2% |
![]()
The two positive numbers are too small to mean anything. The negative one is small but real, with roughly a 1 in 2,500 chance of appearing by luck.
This is an observational study, not an experiment. Site owners added the schema; Ahrefs found the change in its crawler records afterwards and compared. The authors name the catch themselves: "Pages that add JSON-LD often change other things at the same time (e.g. links, content, technical fixes). We can't fully separate schema from these kinds of co-occurrences." And every page in it was already being cited, so it says nothing about whether schema helps a page get discovered in the first place.
A separate test looked at what the engines actually ingest. searchVIU, a German technical-SEO firm, built one product page with prices placed eight different ways - visible HTML, JavaScript-rendered, JSON-LD, hidden microdata, hidden RDFa and so on - then asked ChatGPT, Claude, Perplexity, Gemini and Google AI Mode to fetch the page and report the price.
Nothing hidden in markup was retrieved by anything. Across all four hidden-markup placements, none of the five systems returned a price. What they took from the visible page varied a lot: Gemini found four of eight, ChatGPT three, Google AI Mode two, Perplexity one, and Claude found none at all, replying "I couldn't find any prices displayed on the page."
This is one page, tested once, by a vendor. Treat it as a probe rather than a measurement. Its authors are careful about what it shows: JSON-LD "is NOT read by AI chatbots during direct fetch," but "Schema could be used in earlier phases" - which is the live possibility, that markup still does work through a search index that an engine queries, rather than through the engine reading your page.
There is a plausible mechanism for the direct-fetch result. The open text-extraction pipelines that feed these systems commonly use tools that discard script tags, and JSON-LD lives in a script tag.
The 3.2x Figure You Have Probably Been Quoted
The counter-statistic in circulation is a Princeton and Moz study of 500,000 pages, reporting that FAQ markup makes a page 3.2 times more likely to be cited.
Across this research pass we could not locate it. It is credited to WWW 2026, an academic conference whose published deadlines closed paper submissions on 25 January 2026, while the study is described as running through March 2026. If someone quotes it at you, ask for the paper.
Why Tactics That Win In Tests Lose On Real Sites
Because most were tested on the easier half of the problem, and the half they skip is where they do damage.
An AI answer is assembled in stages. The engine retrieves a set of candidate pages, re-sorts them, then selects what to quote. Nearly every optimisation study you have seen tests only the last stage: researchers hand the model a fixed set of pages and record which one it uses. The Princeton GEO paper that launched this field did that with five supplied documents. So did the 252,000-trial study from Sprinklr researchers, which found formatting edits weak while topical relevance and position in the supplied list dominated.
Both are sound. Neither tells you how to get into the candidate set, which is the part you do not have.
When researchers rebuilt the earlier stages, the tactics reversed. SAGEO Arena, an academic testbed that runs a full search pipeline over 170,000 documents, found that optimising page body text cut a page's chance of being retrieved by about 9%, cut its presence in the top ten after re-sorting by 16%, and cut final citation by 6% - while improving how much a retrieved page gets quoted. You can win the last stage and lose the first two.
And the tactics stop working when everyone uses them. C-SEO Bench, published at NeurIPS 2025, tested a set of optimisation methods across six domains and around 1,900 queries. Counting every method-and-domain pairing, only 3 of 54 produced a real gain, none in question answering, and several pushed pages down. Separately, the researchers varied how many competitors also optimised, and found gains shrinking as adoption rose. They report plain relevance and ordinary ranking work as more effective than any of the tactics.
![]()
The counterweight, and it is a real one. SAGEO Arena also found that structural information helps at the retrieval stage, and the retrieval-systems literature agrees that how a document is divided up affects how well systems fetch it. That is measured inside research pipelines, not in ChatGPT or Google. Structure may well matter through routes that have not been measured in production yet.
How To Read A Statistic In This Field
Three quick checks will disqualify most numbers you are handed.
Does it compare before-and-after, or two different groups? Semrush's content study reports clarity at +32.83%, Q&A format at +25.45%, structured data at +21.60%. Those come from comparing pages AI engines cited against a separate set of pages ranking in Google's top 20. It is a difference between two populations, not a gain from changing anything. Semrush's own method note adds a second problem for a schema article: it says the study "didn't evaluate metadata, HTML structure, schema markup, page layout, or any technical SEO factors" - so whatever its 'structured data' score measures, it is not schema. Both caveats sit away from the numbers, which is where they get repeated from.
Is the outcome citations, or a stand-in for citations? The much-quoted "up to 40% more visibility" from the Princeton GEO paper is Position-Adjusted Word Count: how much of a generated answer reflects a source that was handed to the model. Not citations, not clicks, not traffic.
Does the source still exist? The widely repeated "FAQ schema pages cited at 41% versus 15%" traces to a 50-site study by Relixir, a vendor selling AI-search optimisation. The Internet Archive holds captures of that study page from September 2025 and January 2026. Today the URL and Relixir's entire blog index return 404, while the company site stays up. The number is still circulating.
Two studies also read differently depending on which part you open. SE Ranking's 100,000-prompt analysis summarises that FAQ sections "nearly double your chances of being cited by ChatGPT," while its body reports FAQ pages averaging 3.8 citations against 4.1, and pages carrying FAQ schema averaging 3.6 against 4.2. Its authors flag the gap themselves: the raw comparison is unadjusted, and their model treats missing FAQ sections as a negative once other variables are controlled. OtterlyAI, an AI-search monitoring tool, assigns 52.5% of citations to community platforms in its summary and the same 52.5% to brands in its body.
If you want this done properly, Antonio Blago, an independent researcher, worked through the academic literature against primary sources and excluded vendor blogs by design. He lands where the studies above do.
What Google And Bing Say
Both tell you to write clearly for readers. Neither claims markup gets you cited. That is guidance from the people running the engines, not measurement, and worth having on those terms.
Google published its generative AI optimization guide on 15 May 2026:
"Prioritize effective SEO strategies over 'AEO/GEO hacks': For Google Search, you can ignore tactics like 'chunking' content, creating unnecessary AI text files (like llms.txt), or pursuing inauthentic mentions."
Chunking there means slicing a page into fragments on the theory that models digest them better, a practice we have written about at length. Llms.txt is a plain-text summary file written for AI crawlers. From the same document, the half quoted less often: "Write content for your human audience and make sure the content is well written and easy to follow. People generally appreciate it when web pages are organized by paragraphs and sections, along with headings that provide a clear structure to navigate content."
Microsoft's Bing Webmaster Guidelines say something similar about how Copilot, which runs on Bing, picks sources: one topic per URL, essential information near the top, because "single-topic pages are more likely to be selected for grounding results" - grounding being Bing's word for the sources Copilot pulls in to build an answer.
Google speaks only for Google, and has an interest in what you optimise for. Search Engine Land, the trade publication, ran a piece calling the guidance naive and self-serving, arguing Google "quietly conflates 'we don't use it' with 'you don't need it'." That objection is fair, which is why nothing above rests on Google's word.
One reason to keep schema has quietly expired. Google's hedge is that structured data still earns rich results, the enhanced listings on its results page such as star ratings. Google's documentation changelog records FAQ rich results ceasing to appear on 7 May 2026, and HowTo rich results were retired in 2023. For those two types, that reason is gone.
What Actually Predicts Being Cited
Citation runs on layers: the engine has to reach your page, the page has to be relevant to the question, and then it competes for selection. Off-site evidence is where the strongest measured signal sits.
Ahrefs measured 75,000 brands against Google AI Overview visibility in 2026. Correlation runs from 0, no relationship, to 1, perfect. In web data like this, above 0.5 is strong and around 0.2 is close to noise.
| Signal | Correlation |
|---|---|
| Branded web mentions | 0.664 |
| Branded anchors (links whose clickable text is your brand) | 0.527 |
| Branded search volume | 0.392 |
| Domain Rating (Ahrefs' 0-100 authority score) | 0.326 |
| Referring domains | 0.295 |
| Backlinks | 0.218 |
Mentions beat links, and it is not close. That ordering is consistent with what we found looking at backlinks and AI visibility separately. The top quarter of brands by web mentions averaged 169 AI Overview mentions; the next quarter, 14. Note also backlinks, at 0.218, near the bottom.
But authority is not a requirement. In Ahrefs' March 2026 analysis of 863,000 result pages, only 37.9% of AI Overview citations came from pages in Google's top ten. Treat that one as perishable: Ahrefs ran the same measurement in July 2025 and got 76.1%. It halved in eight months, and there is no reason to think it has stopped moving. About 31% came from pages that do not rank in the top 100 for that query, and 18.2% of those were YouTube. Engines expand a question into sub-questions and fetch sources for those, so a page answering one narrow question precisely can be cited without ranking for the broad query.
None of this is causal, and that matters for what you do with it. Mentions, branded search and links all travel together with brand size, budget and content quality, and we found no published study that separates them. So treat the table as a ranking of where to look first, not a promise. It is still the best-evidenced place to spend, which is a weaker claim than "this causes citations" and the strongest one the evidence supports.
One more: the number of pages on a site correlates with AI visibility at about 0.194. Publishing more is not a strategy.
What To Do
Is the work worth continuing? Yes - the access and content half of it. Stop paying for the markup half, and expect nothing measurable for a quarter.
Drop this week: llms.txt files, FAQPage and HowTo markup added for AI reasons, and cutting pages into machine-sized chunks. Ahrefs scanned 137,210 domains from its own analytics customers, a group it notes skews more technical than the web at large, and found 97% of the llms.txt files it located were never fetched by anything during May 2026.
Then, in order of how well evidenced it is:
- Check AI crawlers can reach you. Everything else depends on it. Our guide to which AI crawlers matter and which are safe to block covers the ones that decide whether you can be cited.
- Put explicit facts on your top pages. Prices, specifications, dates, named numbers. Across 252,000 trials and six models, a missing price was one of the two strongest negative signals, alongside being off-topic. This is the cheapest item here and the best evidenced.
- Find the sub-questions you already half-answer. Export the questions your pages rank for in Search Console, pick the ones where your answer is buried in a paragraph, and give each its own heading with the answer directly beneath.
- Build a mention list, not a link list. Search your category's roundups, comparison posts and directories for competitors' names and not yours. That list is the next two quarters of work, and it is where the strongest correlation in the table points.
- Leave existing schema alone. Removing it gains nothing. Product, Organization and Article still earn rich results.
Measure it against a fixed set of prompts, weekly, for 8 to 12 weeks. AI answers are volatile enough that a single before-and-after check tells you almost nothing. Crawl fixes show up quickly; mention effects take months, if they can be attributed at all.
Before changing anything, it is worth knowing whether AI engines mention you today. CompetLab tracks how ChatGPT, Claude and Gemini describe your brand against your competitors. If you would rather start with crawler access, the AI crawler checker is free and takes a URL.
Frequently Asked Questions
Should I remove the schema markup I already have?
No. Removing it costs time and gains nothing, and Google said when it retired the FAQ and HowTo rich results that structured data which is not being used causes no problems in Search. Keep markup that still earns a rich result, which today means types like Product, Organization and Article. What changes is the spending decision going forward: we found no measured support for adding FAQPage or HowTo markup to win AI citations, and both of those enhanced listings have been retired.
Does llms.txt do anything at all?
Not for AI search visibility, on current evidence. Ahrefs scanned 137,210 domains from its own analytics customers, a group it says skews more technical and SEO-aware than the web at large, and found 97% of the llms.txt files it located were never fetched by anything during May 2026. Google states that Google Search ignores the file. It remains a reasonable convenience for developer tooling and coding assistants, which is how Anthropic and Perplexity use it on their own documentation sites. That is a different job from getting cited in an answer.
If schema doesn't work for AI, why does Google still recommend structured data?
Because it does a different job. Google's wording is that structured data "isn't required for generative AI search" but remains "a good idea to continue using it as part of your overall SEO strategy, as it helps with being eligible for rich results." A rich result is the enhanced listing on Google's results page, such as star ratings. That reasoning holds for schema types with a live Search feature behind them. It no longer holds for FAQPage, whose rich result stopped appearing on 7 May 2026, or for HowTo, retired in 2023. One live possibility is worth keeping in view: markup may still do work through the search index an engine queries, even if the engine does not read it when fetching your page directly.
So is writing FAQ sections a waste of time too?
We found no measurement showing it lifts citations on its own. The one large dataset on it reads both ways: its summary says FAQ sections nearly double ChatGPT citations, while its body reports FAQ pages averaging 3.8 citations against 4.1, and its authors conclude a FAQ section alone will not dramatically change anything. Google and Bing both advise writing clear, single-topic pages that answer a question directly, which is sensible for readers. Treat it as good writing rather than a citation lever.
What is the single most actionable thing here?
Put explicit facts on the page. In a controlled study of 252,000 trials across six models, a missing price was one of the two strongest negative signals, alongside being off-topic. Prices, specifications, dates and named numbers are cheap to add, they are what the engines demonstrably reward, and unlike building brand mentions you can do it this week. It is also the opposite of the common advice to write vaguer, more model-friendly prose.
Can a small site get cited, or is this only for big brands?
It can, and the mechanism is known. In Ahrefs' analysis of 863,000 result pages, about 31% of AI Overview citations came from pages that do not rank in Google's top 100 for that query. Engines expand a question into sub-questions and fetch sources for those, so a page answering one narrow question precisely can be cited without ranking for the broad query. Authority raises the odds without being a requirement, which makes specific, well-answered sub-questions the most realistic route for a small site.
See your AI visibility
Find out how ChatGPT, Claude and Gemini answer when buyers look for tools like yours — where you show up, and against whom.
Share this article