{"id":8609,"date":"2026-08-04T22:06:33","date_gmt":"2026-08-04T22:06:33","guid":{"rendered":"https:\/\/dextora.agency\/?post_type=insight&#038;p=8609"},"modified":"2026-08-04T22:06:33","modified_gmt":"2026-08-04T22:06:33","slug":"ai-citation-factors-ranked-by-evidence-strength","status":"publish","type":"insight","link":"https:\/\/dextora.agency\/en\/insights\/ai-citation-factors-ranked-by-evidence-strength\/","title":{"rendered":"Twenty-Three AI Citation Factors, Sorted by Strength of Evidence"},"content":{"rendered":"<p>Almost everything written about getting cited by AI systems is assertion. Someone ran ten queries, noticed a pattern, and wrote a thread. The useful move is not another experiment but a way of ranking the claims by how much evidence stands behind each one, so that a team can spend its quarter on the top of the list rather than the loudest part of it.<\/p>\n<p>In May 2026 Cyrus Shepard published exactly that: a scored review of <a href=\"https:\/\/signal.zyppy.com\/p\/ai-citation-ranking-factors\" target=\"_blank\" rel=\"noopener\">AI citation ranking factors<\/a>, built by collecting citation studies, experiments, explainers and patents from the previous two years, narrowing to the 54 most useful, and scoring 23 factors on how consistently they showed up.<\/p>\n<p>Before the table, three things about what the scores are, because they get misread immediately.<\/p>\n<ul>\n<li><strong>They measure evidence, not effect size.<\/strong> A high score means many decent studies agree the factor matters. It does not mean the factor moves citations by a large amount.<\/li>\n<li><strong>The scale is not out of ten.<\/strong> The observed range runs from 9.5 down to 2. Nothing scores a 10 and nothing scores a 0, so a low score is the floor of what was measured, not a zero.<\/li>\n<li><strong>It is expert judgement, not statistical pooling.<\/strong> Shepard scored each factor by hand against three criteria: how repeatably a finding appeared across studies, the quality of the underlying data, and whether official documentation or patents supported it. He says plainly that these are not confirmed ranking factors and that correlation is not causation. The 54 sources are not individually listed, so you cannot audit the input set.<\/li>\n<\/ul>\n<h2>The full list, with what it costs and who owns it<\/h2>\n<p>The scores below are Shepard&#8217;s. The cost and ownership columns are ours, added because a ranked list with no price attached is not yet a plan.<\/p>\n<table>\n<thead>\n<tr>\n<th>Factor<\/th>\n<th>Evidence<\/th>\n<th>What it costs to fix<\/th>\n<th>Who owns it<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>URL accessibility<\/td>\n<td>9.5<\/td>\n<td>Hours. A robots or firewall rule<\/td>\n<td>Engineering<\/td>\n<\/tr>\n<tr>\n<td>Search rank<\/td>\n<td>9.4<\/td>\n<td>Quarters. This is the whole SEO programme<\/td>\n<td>SEO and content<\/td>\n<\/tr>\n<tr>\n<td>Fan-out rank<\/td>\n<td>9.3<\/td>\n<td>Months. Cover the sub-questions, not just the query<\/td>\n<td>Content strategy<\/td>\n<\/tr>\n<tr>\n<td>Preview control<\/td>\n<td>9.2<\/td>\n<td>Hours. Snippet directives you may have set years ago<\/td>\n<td>Engineering<\/td>\n<\/tr>\n<tr>\n<td>Query-answer match<\/td>\n<td>9.2<\/td>\n<td>Weeks. Rewrite headings as the questions asked<\/td>\n<td>Editorial<\/td>\n<\/tr>\n<tr>\n<td>Intent-format match<\/td>\n<td>9<\/td>\n<td>Weeks. List queries get lists, comparisons get tables<\/td>\n<td>Editorial<\/td>\n<\/tr>\n<tr>\n<td>Topic cluster ranking<\/td>\n<td>8.9<\/td>\n<td>Quarters. Depth across a subject, not one page<\/td>\n<td>Content strategy<\/td>\n<\/tr>\n<tr>\n<td>Answer near the top<\/td>\n<td>8.8<\/td>\n<td>Days. Move the conclusion above the preamble<\/td>\n<td>Editorial<\/td>\n<\/tr>\n<tr>\n<td>AI-ready structure<\/td>\n<td>8.6<\/td>\n<td>Weeks. Headings, lists and tables that parse cleanly<\/td>\n<td>Editorial and templates<\/td>\n<\/tr>\n<tr>\n<td>Factually specific<\/td>\n<td>8.3<\/td>\n<td>Ongoing. Numbers and names instead of adjectives<\/td>\n<td>Subject expert<\/td>\n<\/tr>\n<tr>\n<td>Explicit phrasing<\/td>\n<td>8.1<\/td>\n<td>Days. Say the thing rather than gesture at it<\/td>\n<td>Editorial<\/td>\n<\/tr>\n<tr>\n<td>Cites sources<\/td>\n<td>8<\/td>\n<td>Days. Link out and attribute<\/td>\n<td>Editorial<\/td>\n<\/tr>\n<tr>\n<td>Self-contained passages<\/td>\n<td>8<\/td>\n<td>Weeks. Each paragraph must survive being lifted alone<\/td>\n<td>Editorial<\/td>\n<\/tr>\n<tr>\n<td>Content visibility<\/td>\n<td>7.6<\/td>\n<td>Weeks to months if the site is client rendered<\/td>\n<td>Engineering<\/td>\n<\/tr>\n<tr>\n<td>Freshness<\/td>\n<td>7<\/td>\n<td>Ongoing. A review cadence, not a bulk date change<\/td>\n<td>Content operations<\/td>\n<\/tr>\n<tr>\n<td>Brand and entity trust<\/td>\n<td>6.8<\/td>\n<td>Years. Cannot be bought quickly<\/td>\n<td>Marketing<\/td>\n<\/tr>\n<tr>\n<td>Length<\/td>\n<td>6.7<\/td>\n<td>Days. Enough to answer, not more<\/td>\n<td>Editorial<\/td>\n<\/tr>\n<tr>\n<td>Language<\/td>\n<td>6.3<\/td>\n<td>Quarters. Real localisation, not machine output<\/td>\n<td>Localisation<\/td>\n<\/tr>\n<tr>\n<td>Entity consistency<\/td>\n<td>5.8<\/td>\n<td>Weeks. Same name, address and description everywhere<\/td>\n<td>Marketing operations<\/td>\n<\/tr>\n<tr>\n<td>Structured data<\/td>\n<td>5.6<\/td>\n<td>Days. Template-level markup<\/td>\n<td>Engineering<\/td>\n<\/tr>\n<tr>\n<td>Known source<\/td>\n<td>5.4<\/td>\n<td>Years. Being a name people already cite<\/td>\n<td>Marketing<\/td>\n<\/tr>\n<tr>\n<td>Domain authority<\/td>\n<td>5<\/td>\n<td>Years. A lagging indicator of everything else<\/td>\n<td>SEO<\/td>\n<\/tr>\n<tr>\n<td>llms.txt<\/td>\n<td>2<\/td>\n<td>An hour, and no evidence it returns anything<\/td>\n<td>Nobody, ideally the build<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>What the shape of that list tells you<\/h2>\n<p>Read the top five together and a pattern appears that is unwelcome to anyone selling a new discipline: <strong>the strongest evidence sits on things that are either plumbing or classic search.<\/strong> Can the page be fetched. Does it rank. Does it cover the sub-questions. Are you allowing a preview at all. Does it answer the query asked.<\/p>\n<p>The first entry is the one most often skipped, because everyone assumes it is fine. Accessibility here means a crawler can actually retrieve the page during grounding, which is not a given if the content only exists after JavaScript or if an edge rule is quietly returning a challenge. That is a ten-minute check and we wrote it up separately as <a href=\"https:\/\/dextora.agency\/en\/insights\/ai-crawlers-javascript-ten-minute-visibility-check\/\">a curl-based visibility test<\/a>. It is worth doing before anything further down this table, because a 9.5 that is broken makes the other twenty-two irrelevant.<\/p>\n<p>The bottom of the list is equally instructive. Domain authority scores 5, and Shepard&#8217;s note is that several studies found a relationship but that it was often weak. llms.txt scores 2 and sits last, with no credible evidence in either direction; we looked at what the crawl data says about that file in <a href=\"https:\/\/dextora.agency\/en\/insights\/llms-txt-does-anything-actually-request-it\/\">a separate piece on whether anything requests it<\/a>. Structured data lands mid-table at 5.6, and the phrasing matters: practically every study that looks at schema finds a positive relationship, the effect is typically small, and it is amazingly consistent. Small and consistent is a good description of a cheap win, which is roughly how we treat it in our <a href=\"https:\/\/dextora.agency\/en\/insights\/schema-org-structured-data-what-to-mark-up-guide\/\">guide to what to mark up<\/a>.<\/p>\n<h2>The overlap number everyone quotes, and how it moved<\/h2>\n<p>The most repeated statistic in this field is that AI Overview citations come disproportionately from pages already ranking. It is true, and the precise version is more interesting than the shorthand.<\/p>\n<p><a href=\"https:\/\/ahrefs.com\/blog\/ai-overview-citations-top-10\" target=\"_blank\" rel=\"noopener\">Ahrefs measured this across 863,000 keyword SERPs and 4 million AI Overview URLs<\/a> and published in March 2026. The figure that circulates as &#8220;38% from the organic top 10&#8221; is actually 37.9% from the first ten SERP blocks, which includes features as well as organic listings. The organic-only number is 37.1%. If you are going to say &#8220;organic&#8221;, use 37.<\/p>\n<table>\n<thead>\n<tr>\n<th>Where the cited URL sat<\/th>\n<th>All SERP blocks<\/th>\n<th>Organic only<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Top 10<\/td>\n<td>37.9%<\/td>\n<td>37.1%<\/td>\n<\/tr>\n<tr>\n<td>Positions 11 to 100<\/td>\n<td>31.2%<\/td>\n<td>26.2%<\/td>\n<\/tr>\n<tr>\n<td>Beyond position 100<\/td>\n<td>31.0%<\/td>\n<td>36.7%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Now the part that rarely travels with the number. Ahrefs states that the same measure was around 76% in July 2025 and roughly 38% by the time of the March 2026 study. The figure halved in about seven months. Ahrefs also notes it improved its parsing methodology between the two studies, and does not quantify how much of the drop is measurement rather than behaviour. So the honest reading is: the overlap is large, it is falling, and part of the fall is instrumentation. Date-stamp it whenever you use it, and do not build a forecast on it.<\/p>\n<p>You will also see a figure of around 90% attributed to seoClarity, which <a href=\"https:\/\/www.seoclarity.net\/research\/aio-rankings-overlap\" target=\"_blank\" rel=\"noopener\">analysed 362,000 queries<\/a>. That is not a contradiction. seoClarity asked whether at least one citation on a given results page also ranks in the top 10, which is a per-page question. Ahrefs asked what share of all citations rank in the top 10, which is a per-citation question. Two different questions, two correct answers. Anyone presenting them as a dispute has not read either.<\/p>\n<p>One genuine divergence is worth knowing: the overlap collapses outside Google. Ahrefs&#8217; <a href=\"https:\/\/ahrefs.com\/blog\/ai-search-overlap\/\" target=\"_blank\" rel=\"noopener\">separate study of standalone AI assistants<\/a> found only about 12% of cited URLs appeared in Google&#8217;s top 10, with Perplexity higher than ChatGPT, Gemini and Copilot. Ranking buys you AI Overviews far more reliably than it buys you chat citations.<\/p>\n<h2>The brand mention claim, with the caveat its authors attached<\/h2>\n<p>The other statistic doing heavy lifting in strategy decks is that brand mentions correlate with AI visibility roughly three times as strongly as backlinks. It comes from <a href=\"https:\/\/ahrefs.com\/blog\/ai-overview-brand-correlation\/\" target=\"_blank\" rel=\"noopener\">Ahrefs&#8217; study of 75,000 brands<\/a>, which found a Spearman correlation of 0.664 for branded web mentions against 0.218 for backlinks. Divide one by the other and you get the three.<\/p>\n<p>Ahrefs&#8217; own conclusion is more restrained than the repost: it describes all the factors studied as showing moderate to very weak correlations, and says correlation is not causation. A 0.664 is a real signal. It is not a mechanism, and it does not tell you that buying mentions produces citations. <a href=\"https:\/\/www.seerinteractive.com\/insights\/what-drives-brand-mentions-in-ai-answers\" target=\"_blank\" rel=\"noopener\">Seer Interactive&#8217;s analysis<\/a> across several hundred thousand keywords points the same way, with Google ranking correlating around 0.65 with mentions in language model answers while backlinks came out weak or neutral.<\/p>\n<h2>How to use the table without wasting a quarter<\/h2>\n<p>The list sorts by evidence, not by return. Sorting by return means crossing it with cost, which is what the third column is for. Done that way, the sequence for most sites is fairly stable:<\/p>\n<ol>\n<li><strong>Verify accessibility and preview control first.<\/strong> Both score above 9, both are configuration, both are frequently broken without anyone noticing.<\/li>\n<li><strong>Fix the editorial mechanics next.<\/strong> Answer near the top, self-contained passages, explicit phrasing, cites sources. Four factors in the 8s, all of them a style guide rather than a project.<\/li>\n<li><strong>Then do the slow, expensive, high-evidence work.<\/strong> Ranking, cluster depth and fan-out coverage are quarters of effort, and they are also the things that were worth doing before any of this existed. Our notes on <a href=\"https:\/\/dextora.agency\/en\/insights\/b2b-content-strategy-writing-that-brings-clients\/\">B2B content strategy<\/a> cover that side, and the <a href=\"https:\/\/dextora.agency\/en\/insights\/technical-seo-checklist-20-points-diy\/\">technical SEO checklist<\/a> covers the plumbing.<\/li>\n<li><strong>Leave the bottom four alone unless they are free.<\/strong> Entity consistency and structured data are cheap enough to do anyway. Domain authority and llms.txt are not things you act on directly.<\/li>\n<\/ol>\n<p>It is worth noting that Google&#8217;s own <a href=\"https:\/\/developers.google.com\/search\/docs\/fundamentals\/ai-optimization-guide\" target=\"_blank\" rel=\"noopener\">guidance on optimising for generative AI features<\/a> arrives at a similar place from the opposite direction: it says there is no separate discipline, no special markup, no requirement to chunk content, and that ordinary SEO practice is the work. When an independent evidence review and the search engine&#8217;s own documentation agree, that is about as much confirmation as this field offers.<\/p>\n<h2>What the list cannot tell you<\/h2>\n<p>It is a synthesis of correlational studies about systems that are probabilistic and change monthly. Every number in it is an average across platforms that behave differently: what earns a citation in AI Overviews is not what earns one in a chat assistant, as the 37% and 12% figures show. Citation volume also varies enormously by platform, and the source mix is heavily skewed toward a handful of large sites, which <a href=\"https:\/\/www.tryprofound.com\/blog\/ai-platform-citation-patterns\" target=\"_blank\" rel=\"noopener\">analysis of hundreds of millions of citations<\/a> has shown consistently.<\/p>\n<p>None of that makes the ranking useless. It makes it a prior, which is more than the field had a year ago.<\/p>\n<h2>The short version<\/h2>\n<p>Twenty-three factors, scored 9.5 down to 2 by how consistently the research supports them, not by how much they move the needle. The top of the list is crawlability, ranking, sub-question coverage and preview settings. The bottom is domain authority and llms.txt. Structured data sits in the middle with a small but remarkably consistent positive effect. The often-quoted top-10 overlap is 37% organic as of March 2026, down from about 76% seven months earlier and partly confounded by a methodology change. And the brand mention correlation is real but described by its own authors as moderate to very weak. Treat the table as a spending order rather than a truth, and start at the top, where the fixes are cheapest and most often already broken.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>The full scored list from Cyrus Shepard&#8217;s review of 54 sources, what each factor costs to fix and who owns it, how the top-10 overlap figure moved, and why brand mention correlations are weaker than they are quoted.<\/p>\n","protected":false},"author":1,"featured_media":8605,"template":"","insight_category":[158],"insight_tag":[174,176,172],"class_list":["post-8609","insight","type-insight","status-publish","has-post-thumbnail","hentry","insight_category-case-studies-2","insight_tag-ai-search","insight_tag-content-strategy","insight_tag-seo"],"acf":[],"_links":{"self":[{"href":"https:\/\/dextora.agency\/en\/wp-json\/wp\/v2\/insight\/8609","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dextora.agency\/en\/wp-json\/wp\/v2\/insight"}],"about":[{"href":"https:\/\/dextora.agency\/en\/wp-json\/wp\/v2\/types\/insight"}],"author":[{"embeddable":true,"href":"https:\/\/dextora.agency\/en\/wp-json\/wp\/v2\/users\/1"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dextora.agency\/en\/wp-json\/wp\/v2\/media\/8605"}],"wp:attachment":[{"href":"https:\/\/dextora.agency\/en\/wp-json\/wp\/v2\/media?parent=8609"}],"wp:term":[{"taxonomy":"insight_category","embeddable":true,"href":"https:\/\/dextora.agency\/en\/wp-json\/wp\/v2\/insight_category?post=8609"},{"taxonomy":"insight_tag","embeddable":true,"href":"https:\/\/dextora.agency\/en\/wp-json\/wp\/v2\/insight_tag?post=8609"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}