AEO fundamentals

Generative engine optimization: what the research actually found.

The term has a birth certificate — one paper, one benchmark, one results table. This page reads it directly: what the authors tested, the numbers as printed, the one they call "up to 40%," the method that scored below the baseline, and what a 2023 study on a prototype does and does not tell you in 2026.

Updated 2026-09-19

Most terms in marketing have no origin. This one has a date. On 16 November 2023, six researchers posted a paper titled GEO: Generative Engine Optimizationto arXiv, coined the phrase “generative engine” for search systems that compose an answer from several sources with a language model, and proposed a practice for getting content into those answers. They built a benchmark to test it, published a results table, and wrote a sentence the whole industry has quoted since: “GEO can boost visibility by up to 40% in generative engine responses.” This page reads the paper rather than the commentary — what was tested, the numbers as printed, and what they do not show. For what the term means next to AEO and who uses which name, that is a different page; this one is the primary source.

The paper, in its own terms

GEO: Generative Engine Optimization. Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande. arXiv:2311.09735. First posted 16 November 2023; latest revision 28 June 2024 (v3); accepted to KDD 2024. The abstract frames the problem from the creator's side: generative engines synthesize answers from many sources, the creators of those sources have little control over whether and how they appear, and GEO is proposed as a “black-box optimization framework” for improving that visibility. Two sentences from the abstract carry the whole claim, and both are worth having exactly:

What was tested

The benchmark. GEO-bench: 10,000 queries, split 8,000 train / 1,000 validation / 1,000 test, drawn from nine sources — MS MARCO, ORCAS-1, Natural Questions, AllSouls, LIMA, Davinci-Debate, Perplexity.ai Discover, ELI5, and GPT-4-generated queries — across 25 domains, tagged along seven dimensions such as difficulty, intent and answer type. For each query the benchmark holds the web sources a generative engine would read to answer it.

The engine. Primarily the authors' own generative engine, built on GPT-3.5-turbo over the top five Google results for each query. Then, as a check against a system the public could use, a second test on Perplexity.ai using 200 test queries.

The measure. Two visibility metrics. The first is a position-adjusted word count — how much of the generated answer is attributable to the modified source, weighted toward earlier placement. The second, which the authors call subjective impression, is a model-scored judgment of how prominently and usefully the source appears. Both are the columns in the table below.

The nine methods.Each takes a source page and rewrites it one way, then measures whether the engine draws on it more. By the paper's own names:

Quotation Addition

Method

Quotation Addition

What it does to the page

add quotations from relevant sources

Position-adjusted word count

27.8

Subjective impression

24.7

Statistics Addition

Method

Statistics Addition

What it does to the page

add quantitative statistics in place of qualitative claims

Position-adjusted word count

25.9

Subjective impression

23.7

Cite Sources

Method

Cite Sources

What it does to the page

add citations to credible sources

Position-adjusted word count

24.9

Subjective impression

21.9

Authoritative

Method

Authoritative

What it does to the page

rewrite in a more authoritative, persuasive tone

Position-adjusted word count

—

Subjective impression

—

Fluency Optimization

Method

Fluency Optimization

What it does to the page

improve fluency and flow

Position-adjusted word count

—

Subjective impression

—

Easy-to-Understand

Method

Easy-to-Understand

What it does to the page

simplify the language

Position-adjusted word count

—

Subjective impression

—

Technical Terms

Method

Technical Terms

What it does to the page

add domain-specific technical terms

Position-adjusted word count

—

Subjective impression

—

Unique Words

Method

Unique Words

What it does to the page

add rare or distinctive vocabulary

Position-adjusted word count

—

Subjective impression

—

Keyword Stuffing

Method

Keyword Stuffing

What it does to the page

add more of the query's keywords to the text

Position-adjusted word count

17.8

Subjective impression

20.2

Numbers are printed exactly as they appear in the paper's main results table on GEO-bench (arXiv:2311.09735, v3). Rows marked "—" have values in the paper that this page does not quote; the table in the paper is the record. We do not convert any of these into our own percentages.

The three that won, and the one that lost

On the first metric the unmodified baseline scores 19.3. Quotation Addition scores 27.8, Statistics Addition 25.9, Cite Sources 24.9. Keyword Stuffing scores 17.8 — below the baseline. On the second metric the order among the top three holds, at 24.7, 23.7 and 21.9, with Keyword Stuffing at 20.2. The reading the industry settled on is the honest one: pages that carry evidence a reader could check — a quotation from a real source, a real figure, a citation — were drawn into synthesized answers more; a page that repeated the query at itself was drawn in less, and on the first measure, less than doing nothing.

The Perplexity check pointed the same way. On a second test on Perplexity.ai using 200 test queries, Quotation Addition scored 29.1 on the word-count metric and Statistics Addition 33.9 on subjective impression — the two methods that led the benchmark led the public engine too.

What it does not prove — our reading, not the authors'

The paper is careful, and the commentary around it mostly is not. Four limits, stated by us:

  • It is a 2023–2024 study on a prototype.The primary engine was built on GPT-3.5-turbo over five Google results; the public check was 200 queries on one engine. The assistants people use in 2026 retrieve differently, use different models, and change often. The authors say this themselves: “While we rigorously test our proposed methods on two generative engines, including a publicly available one, methods may need to adapt over time as GEs evolve.”
  • “Up to 40%” is a ceiling across methods and domains, not a promise per page.The abstract's second sentence — that efficacy varies across domains — is the one commentary drops, and it is the one that matters for a business in a specific vertical.
  • It measured inclusion, not outcomes. More of the answer drawn from your page is not the same as a customer, and nothing in the paper connects the two. Nobody honest promises a citation by a date, and this page does not.
  • “Add statistics” is not “invent statistics.” The method tested replaced vague claims with quantitative ones drawn from the material. A synthesized answer quotes a figure with your name attached, and an invented figure becomes a public correction with your name attached.

How to apply it — the three methods that won, on a real page

This is the section that answers “how do I do generative engine optimization,” and it is short because the paper's result is short. Take a page that answers a real question a buyer asks, and:

  • Replace adjectives with figures that have a named source.Not “most homeowners now use AI” but the survey, the year, and the sample. If you cannot name the source, do not state the number.
  • Quote the authority instead of paraphrasing it. A sentence from the standard, the regulator, the study, the manufacturer — in quotation marks, attributed, in the text where a machine can read it.
  • Cite where a claim came from, inline. A reference the reader could check is a reference the engine can weigh.
  • Do not repeat the query at the page.The one method that scored below the baseline is the one most “SEO” habits reach for first.
  • Then measure. A method that worked on a research benchmark is a hypothesis about your page until the same question, asked again, comes back different.
It depends

Everything above is evidence-shaped content, and it only works on a page that is visible and about a real question. Structured data does not substitute for it — Google states no special markup is required to appear in AI Overviews or AI Mode — and content served only to bots does not either; anything you show a crawler has to be what a person sees. The paper is the evidence. The seven-step playbook is the order of operations.

Where this sits next to everything else on this site

  • AEO vs GEO — what the two names mean, who uses which, and how to buy the work without buying the vocabulary. This page is the primary source behind that one.
  • What is AEO — the plain definition, and how engines decide who to name.
  • How to do AEO — the seven steps and the 25-point checklist; the three winning methods above live inside steps three and four.
  • How to get mentioned by AI — the owner's version, ordered by time to impact.

08 · FAQ

The questions people ask about the paper.

What is generative engine optimization, in one sentence?
The term comes from a single paper — "GEO: Generative Engine Optimization" (arXiv:2311.09735), first posted 16 November 2023 and accepted to KDD 2024 — which coined "generative engines" for search systems that synthesize an answer from multiple sources with a language model, and "GEO" for the practice of changing web content so that it is more likely to be included in those synthesized answers. The paper's own headline claim, quoted exactly, is that "GEO can boost visibility by up to 40% in generative engine responses." Everything else the industry has attached to the term since is commentary on that paper, including this page.
What did the paper actually test?
Nine content-modification methods, applied to the source pages a generative engine reads, measured on a benchmark the authors built: GEO-bench, 10,000 queries drawn from nine sources (MS MARCO, ORCAS-1, Natural Questions, AllSouls, LIMA, Davinci-Debate, Perplexity.ai Discover, ELI5, and GPT-4-generated queries) across 25 domains. The engine was the authors' own generative engine, built on GPT-3.5-turbo over the top five Google results for each query, with a second test on Perplexity.ai using 200 test queries. Visibility was scored two ways — a position-adjusted word-count measure of how much of the answer came from the modified source, and a "subjective impression" measure scored by a model. Those two metrics are the columns in the table on this page.
Which methods worked, and which did not?
As printed in the paper's main results table, on the position-adjusted word-count metric the unmodified baseline scores 19.3; Quotation Addition scores 27.8, Statistics Addition 25.9, and Cite Sources 24.9. Keyword Stuffing scores 17.8 — below the baseline. We print those numbers as the paper prints them rather than converting them into our own percentages; the only percentage we quote is the authors' own "up to 40%." The practical reading is the one the industry has mostly settled on: adding quotations from real sources, real statistics, and citations helped; stuffing keywords did not, and on the paper's first metric it hurt.
Is a 2023 study on a GPT-3.5 prototype still relevant in 2026?
Partly, and the paper says so itself: "While we rigorously test our proposed methods on two generative engines, including a publicly available one, methods may need to adapt over time as GEs evolve." The engines people actually use now are different systems, with different retrieval and different models, and the study was English-language on a research prototype plus a 200-query check on Perplexity. What has held up is the direction rather than the numbers: content that carries real evidence — named sources, real figures, quotable statements — keeps getting drawn into synthesized answers, and content that repeats the query at itself does not. Treat the paper as the origin and the strongest early evidence, not as a rulebook for a specific engine.
Does adding statistics mean inventing them?
No, and this is where the paper is most often misread. The method the authors tested replaced vague claims with quantitative ones drawn from the material; it was never a license to fabricate. On this site the rule is absolute — every number carries a named source, competitor prices come from one verified dataset, and no page ever states a figure it cannot trace — and the reason is not only ethics. A synthesized answer quotes a statistic with your name attached; an invented one becomes a public correction with your name attached. The paper's finding is that real evidence gets cited. Fake evidence gets found.
So how do I actually do GEO?
Apply the three methods that won and skip the one that lost, on pages that answer a real question. Concretely: replace adjectives with figures that have a named source; quote the actual authority instead of paraphrasing it; cite where a claim came from, in the text, where a machine can read it; and never repeat the query at the page hoping to be matched. Then measure whether an answer changed, because a method that worked on a research benchmark is a hypothesis about your pages until you have checked. The full seven-step playbook, with effort and expected effect per step, is on the how-to-do-AEO page linked below; this page is the evidence it stands on.

The research says evidence gets cited. See whether yours is.

The free audit asks ChatGPT, Gemini, and Perplexity ten buyer questions about your business — 30 live answers in about two minutes, three of them shown word for word with every name highlighted. Just an email, no account, no card.

This is a live test: we ask ChatGPT, Gemini and Perplexity what your customers ask and record who gets recommended. Your results are on screen in about two minutes, and we email you a link to the report when the scan is done. We may also send you our newsletter about AI search; every issue has a one-click unsubscribe.

Free · No account · No card · About two minutes

Free audit first: see exactly where you stand before paying anything. 7-day trial, no card.