Engine Mechanics and the Frontier · 14 min read

How an Engine Decides Which Entity You Are

Declaring an identity and corroborating one are different jobs. Almost everything sold as entity SEO does the first and charges for the second.

5xmore likely that an AI model confuses or misidentifies an SME's brand name than a large company'sSearchable, 165 London businesses, 13,365 AI responses
The short version
  • Entity resolution is decided by evidence from parties with no stake in you, not by properties you write about yourself.
  • Google's own Knowledge Vault talk lists schema.org markup as one of four extractors, then says entity linking still has to happen afterwards.
  • Google's rater guidelines instruct evaluators to trust independent sources over what a site says about itself, and explicitly exclude a company's own social profiles from counting as independent.
  • Ahrefs measured minus 4.6% AI Overview citations across 1,885 pages that added JSON-LD, so schema is a parsing layer and not a resolution lever.
  • AI models confused or misidentified an SME's brand name roughly five times more often than a large company's in a 13,365-response test, and the researchers attribute the gap to how much independent validation exists about a company.

Two different problems are wearing the same word

Searchable tested 165 London businesses and collected 13,365 AI responses across ChatGPT, Perplexity and Gemini. Brand name confusion ran at 4% for the small businesses against 0.7% for the large ones, and the models were five times more likely to confuse or misidentify an SME's brand name than a larger company's. Same engines. Same question types. The measured variable was company size, and the researchers' own reading of the gap is how much independent validation exists about a company.

The short answer

Entity resolution is a corroboration problem, not a declaration problem. Markup on your own site tells an engine what you claim to be. Independent sources tell it which thing you actually are. Every knowledge base that has published its rules ranks the second above the first, and several of them refuse to count your own website as evidence at all.

The word entity is doing two jobs, and collapsing them is why entity work stalls.

Job one is recognition. There is a thing here, it is an organization, it is called Summit. Engines have been competent at this for a decade and it costs you nothing.

Job two is resolution. This Summit is that specific organization and not the credit union, the church, the roofing company two states over, or the producer with 400 monthly listeners. Nearly every entity SEO checklist in circulation solves job one and then prices it as if it solved job two.

That gap is the entire post. What follows is the mechanism, the evidence hierarchy the engines and knowledge bases actually publish, a worked disambiguation, and the honest boundary on what your own markup can and cannot do. It belongs alongside the myths worth retiring first, because entity-as-markup is the most expensive one still standing.

What Google published about entity fusion, and why almost nobody quotes it

The last time Google described its entity pipeline in real detail was the Knowledge Vault talk at KDD in August 2014. It is dated and it was explicitly framed as exploratory research, so treat it as architecture rather than current spec. It is still more specific than anything published since.

Four extraction systems feed the pipeline: natural language text, DOM structure, tables, and what the deck calls webmaster annotations. That last one is your schema.org markup. It is not a special channel. It is one of four extractors whose outputs get combined, and the slide covering it ends with a line that should be printed on every entity SEO invoice: "We still need to do entity linking."

Read that again. You handed the engine a JSON-LD block naming your organization, and the system still has to decide which organization that block refers to. Markup makes your claim machine-legible. It does not resolve it.

The rest of the pipeline is a scoring exercise. Every fact carries confidence and provenance. One slide title states the mechanism outright: confidence of true facts rises given more evidence. The worked example in the deck starts with a web extraction confidence of 0.14 for a single fact and lands at a fused belief of 0.61 once graph priors are added. Nothing in that process is a switch you flip. It is a probability you accumulate.

Google's current commercial product tells the same story in plainer language. The Cloud Enterprise Knowledge Graph documentation says the reconciliation confidence score is a metric for the confidence level of the assignment of an entity to a cluster, that confidence falls the further an entity sits from the rest of its cluster, and that scores are bucketed into 0.1 intervals so you should not trust the exact values. The overview page describes an engine that builds a graph to cluster entities into groups using "any combination of fuzzy text, common relationships, entity types, and its attributes."

That is a cloud product, not Google Search, and I am not going to pretend otherwise. But it is the same company describing the same class of problem with the same vocabulary: clusters, distances, confidence, and no user-supplied override. On the consumer side, Google says knowledge panels are automatically generated from information across the web, and that the facts in them come from sources that compile factual information. Notice which verb never appears in any of it: accept.

Four institutions wrote down the same rule, and it is not the one you were sold

The strongest evidence for the corroboration model is not a study. It is that four organizations with entirely different purposes independently arrived at the same policy, and all four rank your own claims last.

Google's Search Quality Rater Guidelines, the 182 page General Guidelines dated September 11, 2025, are the bluntest. Raters are told to look at what a site says about itself, then instructed: "What do outside, independent sources say about them? When there is disagreement between what the website or content creators say about themselves and what reputable independent sources say, trust the independent sources." The reputation section adds that raters should be skeptical of claims a site makes about itself, and even gives them the search syntax: run the query with -site:yourdomain.com.

The detail that should reorganize your entity work is buried in the same section. Google tells raters that a company's own official social media pages "would not be considered independent sources of reputation information about the company." Every LinkedIn, YouTube and X profile in your sameAs chain sits on the wrong side of that line.

Wikipedia's general notability guideline is stricter and older. A topic qualifies when it has received significant coverage in reliable sources that are independent of the subject, and independence "excludes works produced by the article's subject or someone affiliated with it. For example, advertising, press releases, autobiographies, and the subject's website are not considered independent."

Wikidata is the loosest of the four and still requires evidence. Its second notability criterion admits an item that "refers to an instance of a clearly identifiable conceptual or material entity that can be described using serious and publicly available references." Describable using outside references. Not declarable by the subject.

Google Business Profile closes the loop for local. The guidelines ask you to represent your business as it is consistently represented and recognized in the real world across signage, stationery and other branding, and then Google verifies against that world rather than against your form submission.

Four bodies, four purposes, one rule. Rank them by how much they matter to you and the rater guidelines win, not because raters change rankings (they do not, directly), but because that document is the closest public statement of what Google is trying to build systems to approximate. If the target is "trust the independent sources," then the lever is producing independent sources.

A worked disambiguation: Phoenix

Wikipedia's disambiguation page for Phoenix opens with a resolution decision already made. Under the heading Phoenix most commonly refers to, the page lists exactly two entries: the immortal bird of ancient Greek mythology, and Phoenix, Arizona, the capital and most populous city of the state. Everything else sits below that line. Places, dozens of arts and entertainment entries, dozens of businesses, military units, ships, schools, sports franchises and science.

Pick three that a commercial query could plausibly mean.

  • Phoenix, the French alternative rock band.
  • Phoenix Contact, the German industrial connector manufacturer.
  • Phoenix, Arizona, the city.

Now ask what actually separates them for a retrieval system, because it is none of the things sold as entity SEO.

The name string separates nothing. All three are Phoenix. Type helps but is not sufficient, because there is a Romanian rock band called Phoenix too. What resolves them is the shape of the evidence around each one: which other entities they co-occur with, in text written by parties with no stake in the outcome. The band appears next to album titles, festival lineups, producers and chart positions in music press. The manufacturer appears next to product categories, industrial standards, trade registries and distributor catalogues. The city appears next to county names, population figures, census records and neighbouring municipalities.

None of those three won resolution by declaring harder. Each one is legible because a large number of independent parties described it in a consistent context, and the contexts do not overlap.

Now compress that into the case you actually have. A client shares a name with a credit union, a church and a touring act. The credit union has a state charter number, an NCUA record, and local business press. The church has a denominational directory listing and a city permit trail. The act has a MusicBrainz entry, festival listings and reviews. Your client has a website, a Facebook page, and a JSON-LD block with a sameAs array pointing at both of those.

Three of the four Phoenixes are described by outsiders. One is described only by itself. The engine is not being unfair. It is doing exactly what its documentation says it does. This is the mechanism sitting underneath a proper brand accuracy audit, and it is why fixing the eight fields on your own site rarely fixes the answer.

What sameAs actually does, and what it does not

Here is the sentence the industry runs on. Schema.org defines sameAs as the "URL of a reference Web page that unambiguously indicates the item's identity." Unambiguously. Identity. If you only read that, you would reasonably conclude the property settles the question.

Now read what Google says the same property does in its own Organization structured data documentation: "The URL of a page on another website with additional information about your organization, if applicable." Additional information. If applicable. That is not an identity assertion, it is a footnote.

The gap between those two sentences is where most entity SEO budget goes to die.

I sold sameAs chains. For two years I priced entity work as a markup deliverable: twelve to fifteen verified profile URLs, an Organization block, a Person block, done. The output was tidy and clients liked it. What I could never produce, when pushed, was an instance where the chain moved anything that was not already moving. That is why my AI visibility consulting now leads with earned description rather than markup deliverables.

The controlled evidence agrees. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against 4,000 matched control pages and measured minus 4.6% in AI Overviews, plus 2.4% in AI Mode and plus 2.2% in ChatGPT, with the two positives statistically indistinguishable from zero.

Adding schema produced no major uplift in citations on any platform.

Louise LinehanContent Marketer, Ahrefs

Google says the same thing from the other direction. Its AI features documentation states there is no special schema.org structured data that you need to add to appear in AI Overviews or AI Mode. And its structured data policies require that you not mark up content invisible to readers, and prohibit using markup to misrepresent your ownership, affiliation or primary purpose, which is a policy written on the assumption that markup is a claim that can be false.

So keep the markup. It is cheap, it is unambiguous to parse, and a wrong or missing Organization block is a self-inflicted wound. Just stop calling it entity strategy. It is the parsing layer, not the citation lever, and treating a claim about yourself as evidence about yourself is a category error the engines do not make.

Access, measure, map, earn, prove. The same five stages, applied to whether an engine can tell your brand apart from the other four with your name.

See the Cited Method

The corroboration ladder, ranked

Five classes of evidence, ordered by how much weight they carry and how hard they are to fake. The ordering is mine, built from the four policy documents above plus the measured studies, and I will defend the top and the bottom harder than the middle.

Stacked layers showing five classes of entity evidence, from public registries at the top down to self declared markup at the bottom
The layer most entity SEO packages sell is the bottom one. It is the only layer you fully control, which is exactly why it counts for the least.Sources: Google Search Quality Rater Guidelines (September 11, 2025), Wikipedia:Notability, Wikidata:Notability, Google Business Profile guidelines, Ahrefs schema study (May 2026).
Use this graphic on your site

Free to republish with a link back to this page. Copy the embed code:

<a href="https://josephtimpson.com/insights/entity-seo"><img src="https://josephtimpson.com/assets/infographics/entity-seo.svg" alt="Stacked layers showing five classes of entity evidence, from public registries at the top down to self declared markup at the bottom" width="1200" style="max-width:100%;height:auto"></a><p>Graphic by <a href="https://josephtimpson.com/insights/entity-seo">Joseph Timpson</a></p>
Evidence classWhy it carries weightHow much you control it
Public registries and licensing bodiesVerified by a party with statutory duties and no incentive to flatter youAlmost none, which is the point
Independent editorial and sourced reference worksWikipedia's own rule bars the subject's site, press releases and advertising from countingInfluence only, through earned coverage
Verified third-party databases and review platformsGoogle checks the real world representation before it publishes yoursPartial, and verification is the checkpoint
Profiles you control on other platformsConsistency value only, and Google's raters are told these are not independentHigh, which is why it is worth less
Your own site, schema and about pageA claim. Useful for parsing, weak as proofTotal

Two ranking calls worth arguing with. First, I put controlled third-party profiles above your own site even though the rater guidelines call both non-independent, because a LinkedIn or YouTube profile carries a platform-level identity check your footer does not. Second, I put verified databases above controlled profiles rather than alongside them, because verification introduces an outside party into the loop.

The consumer data points the same way. A Skyword survey fielded by Dynata in April 2026 with 1,000 US adults found that when an AI answer conflicts with a brand's own messaging, 29% trust the brand, 12% trust the AI, and 54% go looking outside both to compare.

Consumers are using AI to make real decisions, but when the information feels incomplete or inconsistent, they are looking for proof beyond the brand's own claims.

Andrew WheelerCEO, Skyword

Note what that 54% is measuring. It is not whether people saw your name, it is whether they believed it, which is the same distinction that separates a mention from a citation and makes being named and being cited two different events. That is a vendor-commissioned survey and should be weighted accordingly. It is worth quoting because the behaviour it describes is the same behaviour the rater guidelines instruct, which is a nice piece of convergence: humans and evaluation systems both route around self-description when the stakes are real.

The build order I would actually run

Sequence matters more than completeness here. Doing step four before step one changes nothing, which is the same failure mode as running content before confirming crawler access.

Entity resolution, in the order that works
Step 01

Fix the ground truth

Registry name, address, license numbers, incorporation records and any regulator listing. These are the only records an outsider can check adversarially. If your legal name and trading name diverge, decide which one you are and align the rest to it.

Step 02

Fix what you control, once

Organization markup, an about page a human can verify, consistent name and address, and a sameAs array pointing only at profiles you can prove you own. Budget a day, not a quarter.

Step 03

Force a verification checkpoint

Verified Google Business Profile, industry directories that actually check credentials, professional association listings. Each one puts an outside party between your claim and the record.

Step 04

Earn independent description

Trade press, local news, podcast appearances, conference listings, third-party comparison pages. This is slow and it is the only part that changes the answer, so it gets the budget.

Step 05

Measure the answer, not the markup

Track what engines say about you across a fixed prompt set over repeated runs, and treat any single run as noise.

The fourth step is where the money goes, which is exactly why it gets cut first. It is the same work as what actually earns an AI citation, pointed at your identity instead of at a topic, and you audit its results by reading which sources the engines chose. It is also the step the correlational evidence keeps pointing at: across 75,000 brands, Ahrefs measured branded web mentions correlating with AI Overview visibility at 0.664 against 0.218 for backlinks, with the authors stating plainly that correlation is not causation. Those mentions are the same artefacts that resolve your entity. Earn them. Never buy them, and read the broader case on brand mentions before you build a program around that number.

Where this breaks, and what I cannot prove

The candid section, because the alternative is selling certainty I do not have.

No published evidence links a Wikidata entry to AI citations. I looked hard. Wikidata's scale is real and verifiable, 122.6 million items maintained by roughly 40,700 active editors, and it is a defensible place to be described. But there is no study measuring what an entry does to your citation rate, and anyone quoting one to you is quoting something they did not read.

The Knowledge Vault architecture is twelve years old. Google called it exploratory research at the time. I am using it because it is the most detailed public description of entity fusion that exists, not because I think the 2026 pipeline looks identical.

Cloud Enterprise Knowledge Graph is not Google Search. The confidence-score documentation is the clearest public writing on entity clustering I could find from Google, and it describes a product you buy, not the ranking stack. Same vocabulary, different system.

Measurement of any of this is noisy. SparkToro and Gumshoe ran 2,961 prompts across ChatGPT, Claude and Google AI and found less than a 1 in 100 chance of the same brand list appearing in any two runs of the same prompt. Before and after screenshots of an entity fix are not evidence. Repeated sampling across a fixed prompt set is the minimum bar, which is why sample size discipline is a prerequisite for reporting any of this to a client.

The Searchable study is vendor research. 165 businesses, one city, published by a company selling AI visibility tooling. The direction is consistent with everything else here and the sample is small and geographically narrow. Use it as a signal, not a benchmark.

AI tends to surface and correctly represent companies more often when they have a larger digital footprint, and when they are actively being validated in authoritative parts of the web, such as news publications, directories, review sites and relevant community platforms.

Chris DonnellyCo-founder, Searchable

Minimizing brand drift to ensure accuracy and consistency across every digital touchpoint is now the starting point for securing visibility.

Rachel ThorntonCMO, Adobe Enterprise

The practical test is small enough to run this week. Search your brand name with -site:yourdomain.com appended, the way Google tells its own raters to. Whatever comes back is the evidence an engine has to work with. If the first page is thin, or belongs to somebody else with your name, that is not a schema problem and no property is going to fix it. Fix it the way the rest of the method says to: prove the access, measure the baseline, map the gap, then go earn the description. Unglamorous, slow, and the only part that has ever moved an answer for me, which is why that is the work I do.

Frequently asked questions

What is entity SEO, in practical terms?

It is the work of making a search or AI system resolve your brand to the correct real world thing rather than to something else sharing your name. In practice that means producing verifiable records and independent descriptions, not adding properties to your own markup.

Does adding sameAs to my schema fix entity confusion?

No. Schema.org calls sameAs a URL that unambiguously indicates identity, but Google's own Organization documentation describes it as a page with additional information, if applicable. Keep it because it is cheap and unambiguous to parse. Do not expect it to resolve anything by itself.

Should I create a Wikidata entry for my business?

Only if you can satisfy its criterion honestly, meaning outside references already describe you. No published study measures what a Wikidata entry does to citations or knowledge panels. Treat it as one more place you are described, not as a lever with proven effect.

Why does my client keep getting merged with a similarly named business?

Because the evidence separating them is thin. Engines cluster entities using fuzzy text, shared relationships, types and attributes. When two organizations share a name and neither carries distinct independent coverage, the cluster boundary lands in the wrong place. The fix is producing outside description that only one of you could plausibly attract.

Do knowledge panels come from my structured data?

Not primarily. Google states knowledge panels are automatically generated and that the information comes from sources across the web, with verified entities able to suggest edits. Your markup can inform details like which logo appears. It does not author the panel.

What is the fastest thing I can do this week?

Run your brand name as a search with a minus site operator excluding your own domain, exactly as Google instructs its quality raters. What returns is the evidence an engine has about you. A thin or wrong result set tells you where the real work is.

How do I measure whether entity work is doing anything?

Repeated sampling across a fixed prompt set, never single runs. SparkToro measured under a 1 in 100 chance that two runs of the same prompt return an identical brand list, so a before and after screenshot proves nothing about your entity.

Is local citation building the same as entity corroboration?

Only partly. Directory listings you submit yourself sit closer to controlled profiles than to independent evidence, so bulk citation building adds consistency rather than proof. The ones that carry real weight verify something before publishing, such as licensing bodies, professional associations and platforms that check credentials.

Sources

  1. Google. General Guidelines (Search Quality Rater Guidelines) (2025-09)
  2. Google Machine Intelligence group. Knowledge Vault: a web-scale approach to probabilistic knowledge fusion (KDD 2014 talk) (2014-08)
  3. Google Cloud. Understand the reconciliation confidence score (2026-07)
  4. Google Cloud. Enterprise Knowledge Graph overview (2026-07)
  5. Google. How knowledge panels work (2026-07)
  6. Wikipedia. Wikipedia:Notability (general notability guideline) (2026-07)
  7. Wikidata. Wikidata:Notability (2026-07)
  8. Wikimedia Foundation. Wikidata:Statistics (2026-07)
  9. Google. Guidelines for representing your business on Google (2026-07)
  10. Wikipedia. Phoenix (disambiguation page) (2026-07)
  11. Schema.org. sameAs property (2026-03)
  12. Google Search Central. Organization structured data (2026-04)
  13. Google Search Central. Structured data general guidelines (2026-07)
  14. Google Search Central. AI features and your website (2025-12)
  15. Ahrefs. Does schema markup increase AI citations? (2026-05)
  16. Ahrefs. What correlates with AI Overview brand visibility (2025-05)
  17. The London Business Journal, reporting Searchable research. AI chatbots misinform customers about half of London's small businesses (2026-07)
  18. CMOtech UK, reporting Searchable research. London SMEs suffer AI misinformation in search tests (2026-07)
  19. Skyword (survey by Dynata). When AI Gets Brand Information Wrong, Consumers Look Beyond the Brand (2026-06)
  20. SparkToro with Gumshoe.ai. AIs are highly inconsistent when recommending brands or products (2026-01)
  21. Semrush. Semrush releases expanded 2026 AI Visibility Index (2026-06)
Joseph Timpson
Written by
Joseph Timpson

Joseph Timpson has worked in search since 2010 and runs Timpson Marketing out of St. George, Utah. He built The Cited Method, a five stage framework for earning and proving real citations in AI answers, and publishes what does not work alongside what does.

Want me to run this on your site and show you the before and after?

One call, no pitch deck. We look at what is actually blocking you and tell you the truth about whether we can help.

Book a free consultation