AEO fundamentals
Generative engine optimization: what the research actually found.
The term has a birth certificate — one paper, one benchmark, one results table. This page reads it directly: what the authors tested, the numbers as printed, the one they call "up to 40%," the method that scored below the baseline, and what a 2023 study on a prototype does and does not tell you in 2026.
Updated 2026-09-19
Most terms in marketing have no origin. This one has a date. On 16 November 2023, six researchers posted a paper titled GEO: Generative Engine Optimizationto arXiv, coined the phrase “generative engine” for search systems that compose an answer from several sources with a language model, and proposed a practice for getting content into those answers. They built a benchmark to test it, published a results table, and wrote a sentence the whole industry has quoted since: “GEO can boost visibility by up to 40% in generative engine responses.” This page reads the paper rather than the commentary — what was tested, the numbers as printed, and what they do not show. For what the term means next to AEO and who uses which name, that is a different page; this one is the primary source.
The paper, in its own terms
GEO: Generative Engine Optimization. Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande. arXiv:2311.09735. First posted 16 November 2023; latest revision 28 June 2024 (v3); accepted to KDD 2024. The abstract frames the problem from the creator's side: generative engines synthesize answers from many sources, the creators of those sources have little control over whether and how they appear, and GEO is proposed as a “black-box optimization framework” for improving that visibility. Two sentences from the abstract carry the whole claim, and both are worth having exactly:
What was tested
The benchmark. GEO-bench: 10,000 queries, split 8,000 train / 1,000 validation / 1,000 test, drawn from nine sources — MS MARCO, ORCAS-1, Natural Questions, AllSouls, LIMA, Davinci-Debate, Perplexity.ai Discover, ELI5, and GPT-4-generated queries — across 25 domains, tagged along seven dimensions such as difficulty, intent and answer type. For each query the benchmark holds the web sources a generative engine would read to answer it.
The engine. Primarily the authors' own generative engine, built on GPT-3.5-turbo over the top five Google results for each query. Then, as a check against a system the public could use, a second test on Perplexity.ai using 200 test queries.
The measure. Two visibility metrics. The first is a position-adjusted word count — how much of the generated answer is attributable to the modified source, weighted toward earlier placement. The second, which the authors call subjective impression, is a model-scored judgment of how prominently and usefully the source appears. Both are the columns in the table below.
The nine methods.Each takes a source page and rewrites it one way, then measures whether the engine draws on it more. By the paper's own names:
| Compared on | Methodthe paper's name | What it does to the page | Position-adjusted word countbaseline 19.3 | Subjective impression |
|---|---|---|---|---|
| Quotation Addition | Quotation Addition | add quotations from relevant sources | 27.8 | 24.7 |
| Statistics Addition | Statistics Addition | add quantitative statistics in place of qualitative claims | 25.9 | 23.7 |
| Cite Sources | Cite Sources | add citations to credible sources | 24.9 | 21.9 |
| Authoritative | Authoritative | rewrite in a more authoritative, persuasive tone | — | — |
| Fluency Optimization | Fluency Optimization | improve fluency and flow | — | — |
| Easy-to-Understand | Easy-to-Understand | simplify the language | — | — |
| Technical Terms | Technical Terms | add domain-specific technical terms | — | — |
| Unique Words | Unique Words | add rare or distinctive vocabulary | — | — |
| Keyword Stuffing | Keyword Stuffing | add more of the query's keywords to the text | 17.8 | 20.2 |
Quotation Addition
Method
Quotation Addition
What it does to the page
add quotations from relevant sources
Position-adjusted word count
27.8
Subjective impression
24.7
Statistics Addition
Method
Statistics Addition
What it does to the page
add quantitative statistics in place of qualitative claims
Position-adjusted word count
25.9
Subjective impression
23.7
Cite Sources
Method
Cite Sources
What it does to the page
add citations to credible sources
Position-adjusted word count
24.9
Subjective impression
21.9
Authoritative
Method
Authoritative
What it does to the page
rewrite in a more authoritative, persuasive tone
Position-adjusted word count
—
Subjective impression
—
Fluency Optimization
Method
Fluency Optimization
What it does to the page
improve fluency and flow
Position-adjusted word count
—
Subjective impression
—
Easy-to-Understand
Method
Easy-to-Understand
What it does to the page
simplify the language
Position-adjusted word count
—
Subjective impression
—
Technical Terms
Method
Technical Terms
What it does to the page
add domain-specific technical terms
Position-adjusted word count
—
Subjective impression
—
Unique Words
Method
Unique Words
What it does to the page
add rare or distinctive vocabulary
Position-adjusted word count
—
Subjective impression
—
Keyword Stuffing
Method
Keyword Stuffing
What it does to the page
add more of the query's keywords to the text
Position-adjusted word count
17.8
Subjective impression
20.2
Numbers are printed exactly as they appear in the paper's main results table on GEO-bench (arXiv:2311.09735, v3). Rows marked "—" have values in the paper that this page does not quote; the table in the paper is the record. We do not convert any of these into our own percentages.
The three that won, and the one that lost
On the first metric the unmodified baseline scores 19.3. Quotation Addition scores 27.8, Statistics Addition 25.9, Cite Sources 24.9. Keyword Stuffing scores 17.8 — below the baseline. On the second metric the order among the top three holds, at 24.7, 23.7 and 21.9, with Keyword Stuffing at 20.2. The reading the industry settled on is the honest one: pages that carry evidence a reader could check — a quotation from a real source, a real figure, a citation — were drawn into synthesized answers more; a page that repeated the query at itself was drawn in less, and on the first measure, less than doing nothing.
The Perplexity check pointed the same way. On a second test on Perplexity.ai using 200 test queries, Quotation Addition scored 29.1 on the word-count metric and Statistics Addition 33.9 on subjective impression — the two methods that led the benchmark led the public engine too.
What it does not prove — our reading, not the authors'
The paper is careful, and the commentary around it mostly is not. Four limits, stated by us:
- It is a 2023–2024 study on a prototype.The primary engine was built on GPT-3.5-turbo over five Google results; the public check was 200 queries on one engine. The assistants people use in 2026 retrieve differently, use different models, and change often. The authors say this themselves: “While we rigorously test our proposed methods on two generative engines, including a publicly available one, methods may need to adapt over time as GEs evolve.”
- “Up to 40%” is a ceiling across methods and domains, not a promise per page.The abstract's second sentence — that efficacy varies across domains — is the one commentary drops, and it is the one that matters for a business in a specific vertical.
- It measured inclusion, not outcomes. More of the answer drawn from your page is not the same as a customer, and nothing in the paper connects the two. Nobody honest promises a citation by a date, and this page does not.
- “Add statistics” is not “invent statistics.” The method tested replaced vague claims with quantitative ones drawn from the material. A synthesized answer quotes a figure with your name attached, and an invented figure becomes a public correction with your name attached.
How to apply it — the three methods that won, on a real page
This is the section that answers “how do I do generative engine optimization,” and it is short because the paper's result is short. Take a page that answers a real question a buyer asks, and:
- Replace adjectives with figures that have a named source.Not “most homeowners now use AI” but the survey, the year, and the sample. If you cannot name the source, do not state the number.
- Quote the authority instead of paraphrasing it. A sentence from the standard, the regulator, the study, the manufacturer — in quotation marks, attributed, in the text where a machine can read it.
- Cite where a claim came from, inline. A reference the reader could check is a reference the engine can weigh.
- Do not repeat the query at the page.The one method that scored below the baseline is the one most “SEO” habits reach for first.
- Then measure. A method that worked on a research benchmark is a hypothesis about your page until the same question, asked again, comes back different.
Everything above is evidence-shaped content, and it only works on a page that is visible and about a real question. Structured data does not substitute for it — Google states no special markup is required to appear in AI Overviews or AI Mode — and content served only to bots does not either; anything you show a crawler has to be what a person sees. The paper is the evidence. The seven-step playbook is the order of operations.
Where this sits next to everything else on this site
- AEO vs GEO — what the two names mean, who uses which, and how to buy the work without buying the vocabulary. This page is the primary source behind that one.
- What is AEO — the plain definition, and how engines decide who to name.
- How to do AEO — the seven steps and the 25-point checklist; the three winning methods above live inside steps three and four.
- How to get mentioned by AI — the owner's version, ordered by time to impact.
08 · FAQ
The questions people ask about the paper.
What is generative engine optimization, in one sentence?
What did the paper actually test?
Which methods worked, and which did not?
Is a 2023 study on a GPT-3.5 prototype still relevant in 2026?
Does adding statistics mean inventing them?
So how do I actually do GEO?
Keep reading
AEO fundamentals
AEO vs GEO
The acronym disambiguation page: what GEO means, who uses which term, and why the work is the same.
AEO fundamentals
What is AEO
The plain-English definition, the mechanics, and why the acronym war (AEO vs GEO vs AI SEO) doesn't matter.
AEO fundamentals
How to do AEO
The full practitioner playbook, ordered by effort and impact, with an honest checklist.
AEO fundamentals
SEO vs AEO vs GEO
The three-way question answered once: a ledger, five verdicts, links out for the depth.
The research says evidence gets cited. See whether yours is.
The free audit asks ChatGPT, Gemini, and Perplexity ten buyer questions about your business — 30 live answers in about two minutes, three of them shown word for word with every name highlighted. Just an email, no account, no card.
Free audit first: see exactly where you stand before paying anything. 7-day trial, no card.