Does Schema Causally Lift AI Citations? No.
Schema does three documented jobs and is sold for a fourth. This is a teardown of the primary sources and the three studies, including the one whose headline finding its own author retracted.
- No published study shows that adding schema markup causes an AI engine to cite a page. Three studies with three different designs all land on null or slightly negative.
- Google's own AI optimization guide states there is no special schema.org markup you need to add to appear in its generative AI features.
- Structured data is documented to do three things: help engines parse and disambiguate entities, make a page eligible for one of 31 rich result types, and feed merchant and product surfaces. None of those three is AI citation selection.
- Every uncontrolled study finds schema correlates with citation, because the sites that ship schema are the same sites that already rank, already have entity clarity and already have link authority. Controlling for rank kills the effect.
- The one non-null signal in the literature is attribute-rich Product and Review markup carrying real values like price and rating, not generic Article and Organization wrappers.
The retraction nobody repeated
In February 2026 a study reported that schema markup was negatively associated with AI citation. Odds ratio 0.546, p below .001. Pages carrying schema got cited less often. It was a fantastic headline. It was also wrong, and the person who found it is the one who said so.
Kurt Fischman re-ran the model with query-clustered standard errors and organic rank position in the equation. The negative effect evaporated into nothing. He published the reason himself.
This finding proved to be a methodological artifact: Google's ranking algorithm systematically enriches top-10 organic results for schema-bearing pages, inflating schema prevalence in the non-cited control population.
Read that sentence twice, because it is the whole argument in one line. The pages that rank are the pages that have schema. Any study that measures schema without measuring rank is measuring rank.
No. There is no published evidence that adding schema markup causes an AI engine to cite a page. Three studies with different designs and different samples all land on null or slightly negative. Google states in its own documentation that no special schema.org markup is needed for its AI features. Schema still earns its place in classic search, for entity disambiguation, rich result eligibility and merchant surfaces. Those are three real jobs. AI citation selection is not a fourth.
The argument is always run as all or nothing, which is why it never resolves. One camp says schema is dead because Ahrefs found nothing. The other says schema is more important than ever because Bing said it helps. Both camps are arguing about a claim neither has bothered to state precisely.
So let me state it. There are four separate claims hiding inside "schema helps AI":
- Structured data helps an engine parse a page and resolve which entity it is about.
- Structured data makes a page eligible for a specific enhanced search feature.
- Structured data feeds product, merchant and feed-driven surfaces.
- Structured data increases the probability that a generative engine selects your page as a source for its answer.
Claims one, two and three are documented by the engines themselves. Claim four is the one being sold, and it is the only one with no documentation and no supporting experiment. Separating them is the entire job of this post.
What the engines say in their own words
Start with the party that has the most to gain from you adding markup, because Google's own language is quietly devastating to claim four.
Google's AI optimization guide, last updated 10 July 2026, says it plainly: "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add." The AI features documentation repeats it: "There's also no special schema.org structured data that you need to add."
Google does not stop at neutrality. It tells you why to keep doing it anyway: structured data "helps with being eligible for rich results on Google Search." That is Google endorsing claim two while declining to endorse claim four, in the same paragraph, on the same page.
The single technical prerequisite Google does state for AI features is one almost nobody sells: a page "must be indexed and eligible to be shown in Google Search with a snippet." A stray nosnippet or data-nosnippet zeroes out AI eligibility in a way no amount of JSON-LD can rescue. If you are buying anything mechanical, buy a crawler and snippet access audit before you buy markup.
Now the strongest counter-argument, because it deserves a fair hearing. At SMX Munich in March 2025, Microsoft's Fabrice Canel said something the schema industry has been quoting ever since. Be precise about what the record actually is, because almost nobody repeating it has checked. There is no transcript and no verbatim Canel sentence in circulation. What exists is David Mihm's write-up of the talk, which Barry Schwartz then reported at Search Engine Land: Canel confirmed that schema markup helps Microsoft's LLMs understand your content.
Take it seriously anyway. A Microsoft product manager said it on stage and the trade press stood it up. But look at the verb, because the verb carries the entire load. Understand. Not select, not cite, not rank. It is a claim about parsing, which is claim one. Microsoft's own AI Performance announcement in February 2026, written by four of its AI product managers, gives formatting guidance for AI citation and never once mentions structured data. What it recommends instead is this: "Clear headings, tables, and FAQ sections help surface key information and make content easier for AI systems to reference accurately."
The asymmetry is the finding. When engines describe schema, they describe comprehension. When they describe citation, they describe content. Nobody in the industry seems to have noticed that the two lists do not overlap.
Three studies, three designs, one result
There are now three reasonably serious attempts to measure this, and their value comes from how differently they are built. A single design can be fooled. Three designs failing to find the same effect is harder to dismiss.
| Study | Design | Sample | Result |
|---|---|---|---|
| Ahrefs, May 2026 | Matched difference in differences, 30 days pre and post treatment | 1,885 pages that added JSON-LD, against 4,000 matched controls | -4.6% AI Overviews, +2.4% AI Mode, +2.2% ChatGPT |
| SALT.agency, Sept 2025 | Observational, schema type presence across sites that earned citations | 107,352 websites appearing as citations in Google AI Mode | No type beyond the standard set gave a measurable advantage |
| Fischman, Feb 2026 | Generalized estimating equations, query-clustered errors, rank controlled | 730 AI citations, 75 commercial queries, 1,006 unique pages | Schema presence OR 0.678, p = .296. Rank OR 0.762 per position, p < .001 |
The Ahrefs design is the strongest of the three because it has a treatment and a control and a before and an after. The AI Mode and ChatGPT numbers sit close enough to zero to be noise. The AI Overviews decline is small, statistically significant, and unexplained, which Ahrefs says out loud rather than spinning into a story.
Adding schema produced no major uplift in citations on any platform.
SALT.agency comes at it from the opposite direction. Instead of watching pages add schema, Dan Taylor's team looked at 107,352 sites that had already won an AI Mode citation and asked what markup they carried. If exotic schema types were the edge, the winners would be full of them. They were not.
Schema is a hygiene factor (at best) for AI Mode visibility, not a differentiator.
Fischman's contribution is the one the industry should be reading most closely and is reading least. He is the only one of the three who explicitly modelled organic rank alongside schema, and the result is stark: schema presence is statistically nothing, while every single position of organic rank changes citation odds by roughly 24%. Position one pages were cited in 43% of the queries they appeared in. By position seven that fell to 5%.
One honest caveat on all three. Ahrefs studied pages that already had 100 or more AI Overview citations before treatment, so its data cannot speak to a page with zero existing AI visibility. Fischman's work is a self-published preprint by a practitioner, not peer-reviewed literature. SALT's design is observational and cannot establish direction. None of these is a randomised controlled trial. What they are is three independent teams, using three incompatible methods, all failing to find the thing the industry sells.
The three layers schema is documented to occupy, and the fourth it is sold for
Here is the version of this argument I wish someone had drawn for me two years ago.
Layer one is parsing and entity disambiguation. Google's introduction to structured data says it uses markup "to understand the content of the page, as well as to gather information about the web and the world in general." That second clause is the Knowledge Graph, and it is the most underrated sentence in the documentation. This layer is real, and it is the reason entity clarity work still belongs in a technical scope.
Layer two is rich result eligibility. Google's structured data policies draw the line with unusual precision: "Using structured data enables a feature to be present, it does not guarantee that it will be present." The feature gallery currently lists 31 supported types. Schema.org's core vocabulary contains 823. The other 792 types produce nothing in Google Search, which should tell you how narrow this layer actually is. FAQ and HowTo used to be on that list and are not any more.
Layer three is merchant and feed surfaces. Product markup makes a page "eligible for display in merchant listing experiences on Google Search, including the shopping knowledge panel, Google Images, popular product results, and product snippets," per Google's merchant listing documentation. For ecommerce this is not a nice-to-have, it is the plumbing.
Layer four is AI citation selection. It has no documentation, no stated mechanism, and no supporting experiment. It is drawn detached in the graphic above because that gap is the argument.
There is one more piece of evidence for the appearance-only reading, and it is the sharpest one available. Google says a structured data manual action "means that a page loses eligibility for appearance as a rich result; it doesn't affect how the page ranks in Google web search." Even the penalty for abusing schema is confined to appearance. If markup were a retrieval input, the punishment would touch retrieval. It does not.
The confound, stated cleanly
Every study that finds schema helps is measuring the same three things and calling them one.
The sites that implement structured data are not a random sample of the web. Look at where markup actually clusters, the way the 2025 Web Almanac's SEO chapter tracks it: large publishers, ecommerce platforms, and CMS installs where the SEO plugin emits Organization and Article markup automatically on day one. Those are the same properties that dominate AI citations.
So when a vendor shows you the share of ChatGPT-cited pages carrying structured data, the number may well be accurate and it still proves nothing. Those pages are on managed CMS platforms with editorial teams, established entity footprints, and link authority accumulated over years. The markup and the citation share a cause. Neither caused the other.
Daniel Cheung screened the literature on this question and reached the cleanest formulation of it I have read.
Any schema study that doesn't control for rank is mostly measuring rank.
His review of the schema and citation evidence concluded there is "no credible evidence that schema markup earns AI citations," and he identified the pattern that makes this whole literature legible: the studies that found an effect are the ones missing controls, and the studies with controls found nothing or slightly negative. That split is not a coincidence. It is a diagnostic.
I should be straight about my own position here, because I have been on the wrong side of it. I have built and sold a productized schema service. I still sell it, and I still ship JSON-LD on client sites in volume. What changed is the sentence I am willing to put next to it on a proposal. It used to be about AI visibility. It is now about rich result eligibility and entity disambiguation, because those are the two claims I can defend with a primary source.
If you are being sold schema as an AI visibility lever, or selling it that way, bring the proposal to a call and we will separate the parts that survive the evidence from the parts that do not.
Book a working session→Where the evidence is not null
The intellectually honest version of this post has to include the place the data does not support me, so here it is.
Fischman found one real split. Pages carrying Product or Review markup with concrete attribute values populated, actual prices, actual aggregate ratings, actual specifications, were cited at 61.7% versus 41.6% for pages carrying generic wrappers like empty Article, Organization or BreadcrumbList. That difference was significant at p = .012, and he reported it as most useful to lower-authority domains.
This is worth taking seriously, and it is worth reading carefully, because it does not say what schema vendors will claim it says. The signal is not in the markup. It is in the facts inside the markup. A page with a price, a rating and a spec table has extractable claims on it. A page with an empty Organization block has decoration.
One trap before anyone rushes to add Review markup on the strength of that number: Google's review snippet documentation, updated 24 July 2026, states that if the entity being reviewed controls the reviews about itself, its LocalBusiness or Organization pages are ineligible for the star review feature. The homepage testimonial widget wrapped in AggregateRating is the most common violation in local SEO, and it was already not doing what people thought.
Which means the practical instruction the finding generates is not "add more schema types." It is "put concrete, checkable numbers on the page," and you get most of that benefit from visible on-page content whether or not you wrap it in JSON-LD. That maps directly onto what the original GEO research found in 2023, where adding statistics, quotations and cited sources were the three top-performing interventions, and keyword stuffing scored below doing nothing.
Also worth saying: nobody has good measurement here anyway. AI answers are non-deterministic enough that a single before-and-after screenshot is worthless as evidence, which is why sample size discipline matters more in this discipline than in classic SEO.
What the schema budget buys instead
If you have a fixed number of technical hours and you were going to spend them on markup for AI reasons, here is where the published evidence actually points. Ranked by strength of evidence, not by how good it sounds in a pitch.
| Sold as an AI lever | Supported by evidence | |
|---|---|---|
| Crawler and snippet access | Rarely audited, assumed fine | The one hard prerequisite Google states for AI feature eligibility |
| Schema markup | The headline deliverable | Real for rich results and entities, null for citation across three studies |
| Organic rank on subqueries | Treated as legacy SEO | Strongest measured predictor of citation in the only rank-controlled study |
| Off-site presence | An afterthought | 84% to 93% of citation weight sits on third-party properties in Aleyda Solis's SaaS study |
| Concrete stats and named quotes on page | Content nicety | Top-performing interventions in the original GEO benchmark |
Four things to do with the hours.
First, fix access. Verify that AI crawlers reach your pages and that nothing is suppressing snippet eligibility. This is boring, mechanical, cheap, and it is the only item on this list that Google explicitly states is a requirement.
Second, chase subquery rank rather than head-term rank. Ahrefs' analysis of 1.4 million ChatGPT prompts found cited page titles matched the engine's internally generated fan-out queries more closely (cosine 0.656) than they matched the user's original prompt (0.602). Surfer's fan-out research found pages ranking for both a main query and its fan-outs were 161% more likely to be cited. That is a query fan-out coverage problem, not a markup problem.
Third, earn mentions off your own domain. Ahrefs' 75,000-brand correlation study put branded web mentions at Spearman 0.664 against backlinks at 0.218, and Aleyda Solis measured 84% to 93% of AI citation weight sitting on third-party properties across 15 SaaS brands. Correlational, and both authors say so, but it is a stronger evidence base than schema has. See why brand mentions out-correlate backlinks.
Fourth, put real numbers and real named quotes in the content. It is the one intervention that has both benchmark support and the Fischman attribute-richness result pointing the same direction.
Note what is not on this list: llms.txt, which has its own controlled null result and an explicit statement from Google that it does nothing. The schema argument and the llms.txt argument are the same argument wearing different hats, and they belong together in any honest myths audit.
The verdict, and how to answer the client
Keep the schema. Change the sentence next to it.
Structured data is a legitimate deliverable with three documented jobs, one of which (entity disambiguation) is arguably more valuable in an entity-driven retrieval world than it was five years ago. Microsoft has been reported as saying it aids comprehension. None of that is in dispute and none of it needs defending.
What cannot survive contact with the evidence is the causal claim. When a client or a prospect asks whether schema will get them cited in ChatGPT, the accurate answer has three parts, and it takes about twenty seconds:
- Schema helps engines understand what your page is about and which entity it belongs to. That part is documented by Google and echoed by Microsoft at SMX Munich.
- Schema makes you eligible for specific enhanced search features. Eligible, not guaranteed, and only for the 31 types Google supports.
- Schema has never been shown to cause an AI engine to cite you. Three studies say null, and Google's own docs say no special markup is needed.
- The correlation you have seen quoted is real and uncontrolled. Sites with schema are sites that already rank.
- If you want the citation, buy crawler access, subquery coverage, off-site mentions and concrete on-page facts instead.
Aimee Jurenka, writing at Search Engine Land, put the distinction better than the people arguing either extreme.
This doesn't mean schema is useless, it means schema alone doesn't drive citations.
That is the whole position. Schema is not dead and it is not the mechanism. It never claimed to be the mechanism. The industry claimed that on its behalf, then went looking for numbers to justify the claim, and found correlations it forgot to control.
The strongest evidence in the entire literature is a number about rank, published in a study about schema, by an author who went back and corrected his own headline. Start there. Everything else on the insights index is downstream of getting that ordering right, and if you want the ordering applied to a specific site, that is what the working session is for.
Frequently asked questions
Does schema markup help with AI citations?
No causal evidence exists. Ahrefs tracked 1,885 pages adding JSON-LD against 4,000 controls and found a 4.6% decline in AI Overview citations. SALT.agency found no schema type gave an edge across 107,352 cited sites. Fischman found schema presence statistically null once organic rank was controlled for.
Is schema markup a Google ranking factor?
Google has never claimed it is, and its own policy language implies the opposite. A structured data manual action costs rich result eligibility but explicitly does not affect how the page ranks in web search. If markup were a ranking input, the penalty would touch ranking.
Does Google say you need schema for AI Overviews?
Google says the opposite. Its AI optimization guide states plainly that structured data is not required for generative AI search and there is no special schema.org markup you need to add. The only stated requirement is that a page be indexed and eligible to show with a snippet.
Why do so many AI-cited pages have schema markup then?
Because schema adoption is concentrated in large publishers, ecommerce platforms and CMS installs where SEO plugins emit it automatically. Those are the same sites with editorial resource, entity clarity and link authority. The markup and the citation share a common cause rather than one causing the other.
Didn't Microsoft confirm schema helps its AI?
Search Engine Land reported that Fabrice Canel confirmed at SMX Munich in March 2025 that schema markup helps Microsoft's LLMs understand your content. That is a secondhand account of a talk, and it is a claim about comprehension, not citation selection. Microsoft's own 2026 guidance on being referenced by AI recommends clear headings, tables and FAQ sections, and never mentions structured data.
Is any schema type actually associated with more AI citations?
One split showed up. Fischman found Product and Review markup carrying populated values like price, rating and specifications was associated with a 61.7% citation rate versus 41.6% for generic wrappers, at p = .012. The signal appears to sit in the extractable facts, not in the markup format itself.
Should I remove schema markup from my site?
No. Structured data still drives rich result eligibility across the 31 feature types Google supports, feeds merchant and shopping surfaces, and helps engines resolve entity identity. Those are genuine returns. Remove the AI citation promise from the proposal, not the markup from the site.
What actually predicts whether an AI engine cites a page?
Organic rank is the strongest measured predictor in the only rank-controlled study, at roughly 24% lower citation odds per position. Beyond rank, the evidence points to crawler and snippet access, coverage of fan-out subqueries, off-site brand mentions, and concrete verifiable facts on the page.
Does FAQ schema still do anything in 2026?
Not in Google Search. The FAQ rich result stopped appearing in May 2026 and the documentation was removed in June 2026. FAQPage remains a valid schema.org type and keeping it causes no harm, but it produces no Google feature and no demonstrated AI citation benefit.
Sources
- Ahrefs. We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved. (2026-05)
- SALT.agency. Does schema help you surface more in Google's AI Mode? (2025-09)
- Kurt Fischman. Does Schema Markup Predict AI Citation? A Cross-Platform Empirical Study (2026-02)
- Daniel Cheung. Should You Bother With Schema Markup for AI Search? An Evidence Review (2026-07)
- Google Search Central. Optimizing your website for generative AI features on Google Search (2026-07)
- Google Search Central. AI Features and Your Website (2025-12)
- Google Search Central. Structured Data General Guidelines (2026-07)
- Google Search Central. Structured Data Markup that Google Search Supports (2026-06)
- Google Search Central. Introduction to Structured Data Markup in Google Search (2025-12)
- Google Search Central. Merchant Listing (Product) Structured Data (2026-07)
- Schema.org. Schema.org Vocabulary, Version 30.0 (2026-03)
- Search Engine Land (Barry Schwartz). Microsoft Bing/Copilot use schema for its LLMs (2025-03)
- Search Engine Land (Aimee Jurenka). How schema markup fits into AI search, without the hype (2026-03)
- Microsoft Bing Webmaster Blog. Introducing AI Performance in Bing Webmaster Tools (Public Preview) (2026-02)
- HTTP Archive. Web Almanac 2025, SEO chapter (2026-01)
- Ahrefs. Why ChatGPT Cites the Pages It Cites (2026-04)
- Ahrefs. Branded Web Mentions Correlate with AI Overview Visibility (2025-05)
- Surfer (Joshua Hardwick). The Impact of Query Fan-Out on AI Overview Citations (2025-12)
- Aleyda Solis / Orainti. SaaS AI Search Optimization: Where Citation Weight Actually Sits (2026-07)
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande. GEO: Generative Engine Optimization (KDD 2024) (2024-06)
- Google Search Central. Review Snippet Structured Data (2026-07)
Want me to run this on your site and show you the before and after?
One call, no pitch deck. We look at what is actually blocking you and tell you the truth about whether we can help.
Book a free consultation →