Twenty-Three AI Citation Factors, Sorted by Strength of Evidence

Almost everything written about getting cited by AI systems is assertion. Someone ran ten queries, noticed a pattern, and wrote a thread. The useful move is not another experiment but a way of ranking the claims by how much evidence stands behind each one, so that a team can spend its quarter on the top of the list rather than the loudest part of it.

In May 2026 Cyrus Shepard published exactly that: a scored review of AI citation ranking factors, built by collecting citation studies, experiments, explainers and patents from the previous two years, narrowing to the 54 most useful, and scoring 23 factors on how consistently they showed up.

Before the table, three things about what the scores are, because they get misread immediately.

  • They measure evidence, not effect size. A high score means many decent studies agree the factor matters. It does not mean the factor moves citations by a large amount.
  • The scale is not out of ten. The observed range runs from 9.5 down to 2. Nothing scores a 10 and nothing scores a 0, so a low score is the floor of what was measured, not a zero.
  • It is expert judgement, not statistical pooling. Shepard scored each factor by hand against three criteria: how repeatably a finding appeared across studies, the quality of the underlying data, and whether official documentation or patents supported it. He says plainly that these are not confirmed ranking factors and that correlation is not causation. The 54 sources are not individually listed, so you cannot audit the input set.

The full list, with what it costs and who owns it

The scores below are Shepard’s. The cost and ownership columns are ours, added because a ranked list with no price attached is not yet a plan.

FactorEvidenceWhat it costs to fixWho owns it
URL accessibility9.5Hours. A robots or firewall ruleEngineering
Search rank9.4Quarters. This is the whole SEO programmeSEO and content
Fan-out rank9.3Months. Cover the sub-questions, not just the queryContent strategy
Preview control9.2Hours. Snippet directives you may have set years agoEngineering
Query-answer match9.2Weeks. Rewrite headings as the questions askedEditorial
Intent-format match9Weeks. List queries get lists, comparisons get tablesEditorial
Topic cluster ranking8.9Quarters. Depth across a subject, not one pageContent strategy
Answer near the top8.8Days. Move the conclusion above the preambleEditorial
AI-ready structure8.6Weeks. Headings, lists and tables that parse cleanlyEditorial and templates
Factually specific8.3Ongoing. Numbers and names instead of adjectivesSubject expert
Explicit phrasing8.1Days. Say the thing rather than gesture at itEditorial
Cites sources8Days. Link out and attributeEditorial
Self-contained passages8Weeks. Each paragraph must survive being lifted aloneEditorial
Content visibility7.6Weeks to months if the site is client renderedEngineering
Freshness7Ongoing. A review cadence, not a bulk date changeContent operations
Brand and entity trust6.8Years. Cannot be bought quicklyMarketing
Length6.7Days. Enough to answer, not moreEditorial
Language6.3Quarters. Real localisation, not machine outputLocalisation
Entity consistency5.8Weeks. Same name, address and description everywhereMarketing operations
Structured data5.6Days. Template-level markupEngineering
Known source5.4Years. Being a name people already citeMarketing
Domain authority5Years. A lagging indicator of everything elseSEO
llms.txt2An hour, and no evidence it returns anythingNobody, ideally the build

What the shape of that list tells you

Read the top five together and a pattern appears that is unwelcome to anyone selling a new discipline: the strongest evidence sits on things that are either plumbing or classic search. Can the page be fetched. Does it rank. Does it cover the sub-questions. Are you allowing a preview at all. Does it answer the query asked.

The first entry is the one most often skipped, because everyone assumes it is fine. Accessibility here means a crawler can actually retrieve the page during grounding, which is not a given if the content only exists after JavaScript or if an edge rule is quietly returning a challenge. That is a ten-minute check and we wrote it up separately as a curl-based visibility test. It is worth doing before anything further down this table, because a 9.5 that is broken makes the other twenty-two irrelevant.

The bottom of the list is equally instructive. Domain authority scores 5, and Shepard’s note is that several studies found a relationship but that it was often weak. llms.txt scores 2 and sits last, with no credible evidence in either direction; we looked at what the crawl data says about that file in a separate piece on whether anything requests it. Structured data lands mid-table at 5.6, and the phrasing matters: practically every study that looks at schema finds a positive relationship, the effect is typically small, and it is amazingly consistent. Small and consistent is a good description of a cheap win, which is roughly how we treat it in our guide to what to mark up.

The overlap number everyone quotes, and how it moved

The most repeated statistic in this field is that AI Overview citations come disproportionately from pages already ranking. It is true, and the precise version is more interesting than the shorthand.

Ahrefs measured this across 863,000 keyword SERPs and 4 million AI Overview URLs and published in March 2026. The figure that circulates as “38% from the organic top 10” is actually 37.9% from the first ten SERP blocks, which includes features as well as organic listings. The organic-only number is 37.1%. If you are going to say “organic”, use 37.

Where the cited URL satAll SERP blocksOrganic only
Top 1037.9%37.1%
Positions 11 to 10031.2%26.2%
Beyond position 10031.0%36.7%

Now the part that rarely travels with the number. Ahrefs states that the same measure was around 76% in July 2025 and roughly 38% by the time of the March 2026 study. The figure halved in about seven months. Ahrefs also notes it improved its parsing methodology between the two studies, and does not quantify how much of the drop is measurement rather than behaviour. So the honest reading is: the overlap is large, it is falling, and part of the fall is instrumentation. Date-stamp it whenever you use it, and do not build a forecast on it.

You will also see a figure of around 90% attributed to seoClarity, which analysed 362,000 queries. That is not a contradiction. seoClarity asked whether at least one citation on a given results page also ranks in the top 10, which is a per-page question. Ahrefs asked what share of all citations rank in the top 10, which is a per-citation question. Two different questions, two correct answers. Anyone presenting them as a dispute has not read either.

One genuine divergence is worth knowing: the overlap collapses outside Google. Ahrefs’ separate study of standalone AI assistants found only about 12% of cited URLs appeared in Google’s top 10, with Perplexity higher than ChatGPT, Gemini and Copilot. Ranking buys you AI Overviews far more reliably than it buys you chat citations.

The brand mention claim, with the caveat its authors attached

The other statistic doing heavy lifting in strategy decks is that brand mentions correlate with AI visibility roughly three times as strongly as backlinks. It comes from Ahrefs’ study of 75,000 brands, which found a Spearman correlation of 0.664 for branded web mentions against 0.218 for backlinks. Divide one by the other and you get the three.

Ahrefs’ own conclusion is more restrained than the repost: it describes all the factors studied as showing moderate to very weak correlations, and says correlation is not causation. A 0.664 is a real signal. It is not a mechanism, and it does not tell you that buying mentions produces citations. Seer Interactive’s analysis across several hundred thousand keywords points the same way, with Google ranking correlating around 0.65 with mentions in language model answers while backlinks came out weak or neutral.

How to use the table without wasting a quarter

The list sorts by evidence, not by return. Sorting by return means crossing it with cost, which is what the third column is for. Done that way, the sequence for most sites is fairly stable:

  1. Verify accessibility and preview control first. Both score above 9, both are configuration, both are frequently broken without anyone noticing.
  2. Fix the editorial mechanics next. Answer near the top, self-contained passages, explicit phrasing, cites sources. Four factors in the 8s, all of them a style guide rather than a project.
  3. Then do the slow, expensive, high-evidence work. Ranking, cluster depth and fan-out coverage are quarters of effort, and they are also the things that were worth doing before any of this existed. Our notes on B2B content strategy cover that side, and the technical SEO checklist covers the plumbing.
  4. Leave the bottom four alone unless they are free. Entity consistency and structured data are cheap enough to do anyway. Domain authority and llms.txt are not things you act on directly.

It is worth noting that Google’s own guidance on optimising for generative AI features arrives at a similar place from the opposite direction: it says there is no separate discipline, no special markup, no requirement to chunk content, and that ordinary SEO practice is the work. When an independent evidence review and the search engine’s own documentation agree, that is about as much confirmation as this field offers.

What the list cannot tell you

It is a synthesis of correlational studies about systems that are probabilistic and change monthly. Every number in it is an average across platforms that behave differently: what earns a citation in AI Overviews is not what earns one in a chat assistant, as the 37% and 12% figures show. Citation volume also varies enormously by platform, and the source mix is heavily skewed toward a handful of large sites, which analysis of hundreds of millions of citations has shown consistently.

None of that makes the ranking useless. It makes it a prior, which is more than the field had a year ago.

The short version

Twenty-three factors, scored 9.5 down to 2 by how consistently the research supports them, not by how much they move the needle. The top of the list is crawlability, ranking, sub-question coverage and preview settings. The bottom is domain authority and llms.txt. Structured data sits in the middle with a small but remarkably consistent positive effect. The often-quoted top-10 overlap is 37% organic as of March 2026, down from about 76% seven months earlier and partly confounded by a methodology change. And the brand mention correlation is real but described by its own authors as moderate to very weak. Treat the table as a spending order rather than a truth, and start at the top, where the fixes are cheapest and most often already broken.