AI Search Ranking Signals: Verified, Contested, and Unmeasurable

Key Takeaways
- Off-site brand signals have the strongest measured correlations. Branded web mentions score 0.664, and YouTube mentions 0.737 against AI visibility, while backlinks reach only 0.218 (Ahrefs, 75,000 brands).
- Existing search visibility is a genuine predictor. Ranking on Bing's first page has been linked to roughly 3x higher odds of ChatGPT citation.
- Freshness is real but was inflated. Ahrefs measured AI-cited URLs averaging 25.7% newer; the widely repeated "4.3x" version of that finding is unsourced and should be dropped.
- Schema markup correlates but doesn't cause. A controlled test of 1,885 pages that added JSON-LD found no meaningful citation lift.
- E-E-A-T as a correlation coefficient is not measurable. As a set of practices - named authors, credentials, cited sources - it's well supported. Those are different claims.
Why Do Factor Lists Mislead?
Three recurring problems make these lists less useful than their format suggests.
Correlating against unmeasurable concepts. E-E-A-T, topical authority, and semantic relevance are frameworks, not metrics. To correlate them, you must first operationalize them into something countable, and that choice determines the result. When the operationalization isn't published, the coefficient is decorative.
Treating correlation as instruction. Cited pages being three times more likely to carry JSON-LD is a real observation. It became "add schema to get cited," which controlled testing does not support.
🛍️✨ From the author · ByteStoreSEO for Beginners: Get Found on Google₹10,000View offer →Ignoring the two-stage process. Retrieval decides what enters the candidate pool; selection decides what gets cited. A signal can matter enormously at one stage and not at all at the other, which is how two honest studies reach opposite conclusions about the same factor.
Which Signals Hold Up, and Which Do Not?
- Tier A means measured at scale with disclosed methodology, or mechanically necessary.
- Tier B means plausible with supporting observation but no controlled proof.
- Tier C means contested, null under testing, or unmeasurable, as usually stated
| Signal | Tier | Evidence |
|---|---|---|
| Branded web mentions | A | 0.664 correlation across 75,000 brands |
| YouTube mentions | A | 0.737, the strongest single measured signal |
| Existing search visibility | A | ~3x citation odds from Bing page one; #1-ranked pages cited at 43.2% by ChatGPT |
| Crawlability and retrievability | A | Mechanically necessary - an unfetchable page cannot be cited |
| Freshness | A | AI-cited URLs average 25.7% newer than traditionally cited ones |
| Branded search volume | B | 0.334 - 0.392 across independent datasets |
| Topical depth and coverage | B | Consistent with fan-out retrieval; no agreed metric |
| Entity clarity and consistency | B | Plausible mechanism, largely untested at scale |
| Author credentials and sourcing | B | Aligns with rater guidance; not independently isolated |
| Answer-first structure | B | Aids retrieval; 8% overlap at the citation stage |
| Backlinks | B | 0.218 - real but far weaker than brand mentions |
| Domain authority metrics | C | 0.266 - 0.326 in one dataset, ~0.18 in another, negative in some verticals |
| Schema markup | C | 3x correlation, null in a controlled test of 1,885 pages |
| E-E-A-T as a coefficient | C | Not a score; unmeasurable as stated |
| llms.txt | C | ~7-10% adoption; visibility-driving bots rarely fetch it |
What Tier A Actually Tells You?
The three strongest measured signals are all off-site or pre-existing: brand mentions, video presence, and search visibility you already have. None is a page-level tactic.
That's the uncomfortable finding for anyone selling on-page GEO work. In the Ahrefs dataset, backlinks - the load-bearing signal of two decades of SEO - sit at 0.218, below unlinked brand mentions at 0.664 and well below YouTube at 0.737. The distributional effect is stark too: brands in the top quartile for web mentions earn roughly ten times more AI Overview mentions than the next tier.
The plausible mechanism is that language models learn brand associations from text co-occurrence. A link is markup; a mention is language. YouTube fitting at the top is consistent - transcripts are plain text at enormous scale.
What Tier B Deserves?
Tier B isn't dismissal. Topical depth, entity clarity, author credentials, and answer-first structure are all sensible practices with plausible mechanisms. They're just not proven at the level their proponents claim.
Answer-first structure is the clearest case. Fan-out query wording overlapped 53–59% with retrieved results but only 8% with cited source descriptions. Structure helps you get retrieved. It doesn't decide selection.
What Tier C Requires Honesty About?
Schema markup. The correlation is real, and the causal claim isn't supported. Ahrefs' controlled test found no lift; Otterly's 319-prompt experiment found no isolated connection; Fischman's study concluded attribute-rich schema helps only lower-authority domains. Implement it for entity clarity, then stop optimizing it.
Domain authority. Reported figures range from 0.266 - 0.326 down to about 0.18, and negative in some verticals. A signal whose sign flips between studies isn't a reliable lever.
E-E-A-T. This needs care, because the practices are genuinely valuable while the metric is fictional. Named authors with real credentials, cited primary sources, first-hand evidence, and honest limitations all plausibly help - they're what quality raters assess and what makes content credible to a reader. What doesn't exist is an E-E-A-T score to correlate against. Treat published E-E-A-T coefficients as unverifiable and the underlying practices as worth doing anyway.
Which Claims Should Be Retired?
- "Freshness gives a 4.3x advantage." The sourced figure is 25.7% newer on average. The multiplier version is unsourced inflation of a real finding.
- "96% of AI Overview content comes from verified E-E-A-T sources." Verified by whom, against what threshold? Unfalsifiable as stated.
- "Vector embedding alignment drives 7.3x higher selection." No accessible methodology.
- "Schema markup triggers AI citations." Contradicted by controlled testing.
- "Domain authority no longer matters." Overcorrection. It's weak and inconsistent, not zero.
A useful filter: if a statistic has two decimal places and no linked methodology, treat it as marketing until proven otherwise.
Why Do Credible Studies Disagree?
Two apparent contradictions come up constantly, and both have clear explanations.
"92.36% of AI Overviews cite at least one top-10 domain" versus "citations from top-10 pages fell from 76% to 38%." These measure different things. The first asks whether any cited domain also ranks; the second asks what share of citations come from ranking pages. Both can be true, and together they say Google's AI still draws on ranking sites while increasingly citing pages that don't rank.
"Rankings predict citations" versus "88% of cited URLs don't rank in the top 10." Domain versus URL again, plus head keywords versus long-tail prompts. Studies using short keywords report high overlap; studies using conversational prompts report low overlap.
Neither pair is a real disagreement. They're different questions being quoted as if they were the same one.
What Should You Actually Do?
Fix retrieval first, because it's absolute and cheap. Crawler access, rendering, indexation. An unfetchable page scores zero on every other signal.
Fund brand presence, because that's where the correlations are. Earned coverage, original data worth citing, video, category presence. This is the least novel and most supported recommendation in the entire field.
Keep freshness genuine. Update content substantively where recency affects accuracy. Don't game dateModified.
Do the Tier B practices anyway. Author bylines, sourced claims, topical coverage, answer-first sections. Cheap, sensible, plausibly helpful - just don't budget as if they were proven.
Stop buying Tier C as strategy. Validate schema and move on. Skip llms.txt. Ignore anyone quoting an E-E-A-T coefficient.
How to Grade Your Own AI Visibility Roadmap?
Five steps to find where your budget actually sits.
- List every AI visibility activity from the last quarter. One line each: schema work, content refreshes, digital PR, technical fixes, tooling subscriptions.
- Assign each a tier using the table above: A for well-measured signals, B for plausible-but-unproven, C for contested or unmeasurable.
- Attach the cost. Hours, spend, or both. Estimates are fine; precision is not the point.
- Total the cost per tier. If most of your spend sits in B and C, that is the finding, and it is the common result.
- Check every statistic you have quoted to a client or stakeholder against the retired-claims list. Replace anything you cannot trace to an accessible methodology.
Run it once per quarter. The exercise takes under an hour and usually reorders the roadmap by itself.
Conclusion
The honest version of an AI ranking factors list is shorter and less satisfying than the popular version. A handful of signals are well measured, most are plausible but unproven, and several widely repeated numbers don't survive a look at their sourcing.
The strongest evidence points somewhere no optimization checklist wants to go: become a brand that gets mentioned, in text, across the places models read. That's slower than adding schema, and it's the finding the data keeps returning.
Auditing your own program? Take your AI visibility roadmap and grade each item A, B, or C against this table. If most of your budget sits in C, that's the finding.

💬 Comments (0)
Top · NewestSign in to comment