AntCrow
(Guides · Search & AI)

What AI search actually cites

Most AEO advice is recycled guesswork, and several of its most repeated tactics are contradicted by the largest datasets available. This is what the published research supports, what it refutes, and where the honest uncertainty sits.

Reviewed by AntCrowLast reviewed 11 min read

Answer engine optimisation has attracted an enormous amount of confident advice in a very short time, most of it untested. Meanwhile several large-scale studies have been published, and some of what they found directly contradicts the tactics being sold. This page separates the two, and flags where the evidence is genuinely thin rather than pretending otherwise.

The strongest signals are off your website entirely

This is the uncomfortable finding. Ahrefs analysed correlations between brand visibility in AI answers and a range of factors across roughly 75,000 brands, and the strongest correlate was mentions on YouTube, at around 0.74. Branded mentions across the web generally followed at roughly 0.66 to 0.71. Domain Rating, the classic SEO authority metric, sat far behind at around 0.27 to 0.33, and raw backlink counts were weaker still. Correlation is not causation and brand size plausibly drives all of these together, but the ordering is consistent and it points somewhere awkward: the biggest lever on whether an AI recommends you is how much the rest of the web talks about you, not how well your own pages are optimised.

0.74

Correlation between YouTube brand mentions and AI visibility

0.66

Correlation for branded web mentions generally

0.27

Correlation for Domain Rating, the classic SEO metric

Community platforms are cited far above their share of the web

Analyses of AI citation sources repeatedly surface the same names near the top: Reddit, YouTube, LinkedIn, and Wikipedia. SE Ranking's study of roughly 129,000 domains found that sites with a heavy Reddit presence averaged around 7.0 ChatGPT citations against roughly 1.8 for sites with minimal presence. Review platforms showed a similar pattern, with listed businesses averaging several times the citations of absent ones. Exact figures vary substantially between studies because different researchers measure citation share and mention share differently, so treat the specific numbers as indicative. The direction is consistent enough to act on: presence on the platforms where people discuss your category matters more than most on-site work.

Tactics the data does not support

Three pieces of standard AEO advice come out badly in the largest available dataset, and it is worth being blunt about them because they are sold constantly.

  • Question-phrased headings. SE Ranking found pages using standard topical headings averaged more citations than those using question-style headings. This is the single most repeated AEO tactic and the data suggests it is mildly counterproductive. Use question headings where readers genuinely search that way, not as a citation technique.
  • FAQPage schema. The same study found a slight negative association with citations, and Google retired the FAQ rich result entirely in May 2026. The markup does no harm, but adding it expecting AI citations is not supported.
  • Keyword-optimised titles and URLs. Pages with heavily keyword-matched titles averaged materially fewer citations than pages with natural ones. The same pattern appeared for URL slugs. This echoes the Princeton GEO research, which found keyword stuffing was the only tested modification that reduced visibility.

Tactics the data does support

The Princeton GEO paper, published at KDD 2024, tested nine content modifications against a generative engine across roughly 10,000 queries. Citing sources performed best, followed by adding statistics and adding direct quotations. Notably, the effect was largest for pages ranking around fifth rather than first, which suggests this work matters most for the businesses that are currently almost visible rather than already dominant.

  1. Cite your sources with real outbound links. The highest-performing single modification in the Princeton research. Outbound citation is what makes a page read as evidence rather than as assertion.
  2. Put concrete numbers in. Pages carrying many specific data points are cited substantially more than pages of adjectives. Every number needs a source attached, or it becomes a liability rather than an asset.
  3. Front-load the answer. Analysis of ChatGPT citations found a large share came from the first third of a page. Lead each section with its conclusion, then support it.
  4. Write sections that stand alone. Roughly 120 to 180 words between headings performed best, which is notably longer than the 40-word answer blocks commonly recommended. A chunk needs claim, evidence, and enough context to survive being extracted.
  5. Name things explicitly. Heavily cited text carries a much higher density of named entities than ordinary prose. Replace "our platform" and "the solution" with actual names, versions, places, and dates.
  6. Show a visible last-updated date. Pages displaying one were cited substantially more often, and the overwhelming majority of cited content is recent. The date has to be true.
  7. Attribute content to a named person with real credentials, not to a faceless team.
  8. Be fast. Citation rates correlated strongly with first contentful paint. Live-fetch agents time out on slow pages exactly like impatient humans do.

There is no single AI visibility to win

Different engines cite startlingly different sources for the same question. Published comparisons have found that Google's own AI Overviews and AI Mode share only a small minority of cited URLs with each other, and that the large majority of cited URLs appear in exactly one engine. Semrush's analysis of AI prompts also found ChatGPT surfacing many more sources per answer than Gemini does, which makes Gemini a far narrower funnel. The practical consequence is that a single AI visibility score is close to meaningless, and progress has to be measured engine by engine.

The uncomfortable summary

On-site structure work is real, cheap, and bounded. Do it: answer-first sections, real numbers with sources, named authors, visible dates, fast pages, explicit entities. Then accept that the larger share of the outcome sits in things that take much longer, namely being genuinely talked about on the platforms where your category gets discussed. Anyone offering to make you visible in AI search purely through changes to your own website is selling the easy half and quietly omitting the hard one.

Common questions

  • SEO optimises for a position in a ranked list of links. AEO, answer engine optimisation, optimises for being cited inside an AI-generated answer. GEO, generative engine optimisation, is a near-synonym of AEO originating in academic research and used interchangeably in practice. All three share a foundation of crawlable, credible content, but AEO and GEO put more weight on self-contained answers, concrete evidence, and clear entity signals.

Sources

Every factual claim on this page traces to one of the following. If you find something here that is out of date, we would genuinely like to know.

  1. GEO: Generative Engine Optimization. Aggarwal et al., KDD 2024, arXiv:2311.09735.
  2. AI brand visibility correlations study. Ahrefs.
  3. How to optimize for ChatGPT: a data study. SE Ranking.
  4. AI features and your website. Google Search Central.
  5. New AI visibility insights in Bing Webmaster Tools. Microsoft Bing Blogs.
  6. 2026 AI Visibility Index. Semrush.

Want this handled rather than explained?

AntCrow builds websites with all of the above built in. Tell us what you are trying to achieve and we will tell you honestly what is worth doing.

Start a project