Guide

Generative Engine Optimization: What Survived Testing

Generative engine optimization is a named method from a 2024 KDD paper that measured edits to a document already inside an answer engine’s context. Its headline gain describes share of an answer, not discovery. A later benchmark found most of those methods ineffective and often harmful to ranking.

By the Addition team Updated 10 September 2026 12 min read

Where generative engine optimization came from

The definition of GEO is not in dispute. Where the word came from matters more. The paper that introduced it measured something specific, and it measured it in a narrower situation than the number sounds.

The word has a paper behind it.

Generative engine optimization was named in a paper by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande. It was published in the proceedings of the 30th ACM SIGKDD conference on 24 August 2024. The paper is open access and runs from page 5 to page 16.

What that number measures and what it does not

The paper introduced GEO-bench, a benchmark of 10,000 queries. It measured visibility with position-adjusted word count: how much of a generated answer traces back to a given source, weighted so earlier sentences count for more. That is a share-of-answer measure. It is not traffic, not ranking, and not whether the engine found you.

The distinction matters because of how the experiment was built. Five documents were already placed in the model’s context, and the question was which of them the answer would draw from. The consequence is straightforward. A forty percent gain there does not mean forty percent more traffic, forty percent higher ranking or forty percent more clicks. It means the optimised source contributed more to the answer in that pipeline.

Everything about how Google’s own AI surfaces behave, and how far citation has drifted from ranking on them, is a separate question with its own measurements. AI search engine optimization takes those apart. What comes next is the method itself and the testing it has produced since. The shorter definition of its sibling term sits in answer engine optimization, and what the GEO tools sold under this name buy you is a separate page.

How nine tested edits performed

The method the paper proposed is not a philosophy. It is nine concrete edits to a document, each one measured separately, which is why the testing that followed could argue with it. The results are more interesting than the headline, because they point in a consistent direction.

Values from the critical survey’s normalised table, read 3 September 2026. The arithmetic is the survey’s, not ours.
What each edit did to share of the answer

Read the order, not the size.

The survey that reviewed this literature draws two readings out of that ordering, and both are its own. First, directly extractable information, meaning figures, definitions, quotations and references, can facilitate the use of a document. Second, in its words, keyword stuffing does not transfer effectively from conventional SEO. What the ranking on its own establishes is the order, whereas the reason for it stays an interpretation. The order is the useful part. The cheap edits are the ones that add something a reader could check, not the ones that reword what is already there.

A second result rarely quoted

There is a second result inside the same experiment that rarely gets quoted, and it changes what the advice means. When Cite Sources was applied, the fifth source in the context gained 115.1 percent while the first lost 30.3 percent. The metric is a share of one answer, so it is competitive by construction. Whatever the fifth source gained, it gained from somebody.

A short definition of the neighbouring term, and where the two words came from, sits in answer engine optimization.

What a benchmark found across six domains

Two years after the paper, the obvious question had an answer. Haritz Puerto, Martin Gubri, Tommaso Green, Seong Joon Oh and Sangdoo Yun built C-SEO Bench and published it at NeurIPS Datasets and Benchmarks in 2025. It tests the methods across two tasks and three domains each, including product recommendation.

The three findings this page leans on, in one paragraph

The arXiv abstract for C-SEO Bench, stating that most current C-SEO methods are not only largely ineffective but frequently have a negative impact on document ranking, that traditional SEO strategies are significantly more effective, and that gains decrease as adopters increase, giving the problem a congested and zero-sum nature. The metadata line reads Accepted at NeurIPS Datasets and Benchmarks 2025
  1. 1Ineffective is the weaker of the two words. The sentence says frequently negative, so a method can move a page down rather than leave it where it was.
  2. 2Reviewed, and the frame is where you can tell: accepted at NeurIPS Datasets and Benchmarks 2025. Most of the other work cited on this page is a preprint.
  3. 3Two search tasks and three domains each, so the six domains are the breadth of the test and also its limit.
arxiv.org, C-SEO Bench: Does Conversational SEO Work, read 10 September 2026.

On stores the result is the same shape. E-GEO, by Bagga, Farias, Korkotashvili, Peng and Wu, ran 13,747 product queries against five generative engines.

Martinez’s 2026 critical survey reports that ten of E-GEO’s fifteen starting heuristics came out neutral or negative. The count is the survey’s, not the testbed authors’.

Counts from C-SEO Bench by Puerto and colleagues, as tallied in Martinez’s 2026 critical survey, read 3 September 2026.
What survived the benchmark

Two years later, the answer is narrow.

Most current methods are not only largely ineffective but frequently have a negative impact on document ranking, and that is the paper’s own summary of its results. Several transformations reduce rank outright. What matters is the shape of the test behind that sentence: how many methods, in how many domains, and which task produced nothing at all.

The founding number and the benchmark that tested it belong side by side, and the two are not in conflict once you notice they measure different stages. The paper measures how much of an answer a document contributes once it is already in the context. The benchmark measures whether the document gets into the ranking at all. An edit could raise the first and lower the second. That is the shape the two literatures suggest together, though no single experiment measured both on the same edit.

Where the evidence points instead

That leaves a narrow, boring instruction, and it comes from the research. The same benchmark reports that traditional SEO strategies, the ones aiming to improve the ranking of the source in the model’s context, are significantly more effective. Google puts it from the other side. Its own guidance says the best practices for SEO continue to be relevant, because these features rely on core Search ranking systems. The e-commerce testbed lands there too. Across 13,747 product queries, the survey reviewing it reports that ten of the fifteen hand-written heuristics came out neutral or negative. What beat them was optimisation done systematically, not from a list, which is the same conclusion the answer engine optimization entry reaches about checklists.

  1. Get retrievable first

    Google states the eligibility requirement plainly: a page has to be indexed and eligible to be shown with a snippet. Nothing downstream matters until that holds, and on a store the usual blockers are the ones an ecommerce SEO audit looks for.

  2. Put extractable material on the page

    Zhang and colleagues found that pages with high citation influence are longer, more structured and richer in definitions, numerical facts, comparisons and procedural steps. That is the same direction the paper’s three winning edits point in.

  3. Change little, and change the right thing

    The AgentGEO work reports over 40 percent relative improvement in citation rates while modifying only 5 percent of a document, against 25 percent for baseline methods. The same paper warns that generic optimisation can harm long-tail content.

  4. Stop before the checklist

    The standard GEO list is where the evidence runs out, and some items on it have been called unnecessary by the platforms themselves.

If you want crawler access, indexing and structure checked instead of argued about, an ecommerce SEO audit covers exactly that.

What the platforms call unneeded

Crawler access, indexing and structure are the parts you can have checked. Past them sit two kinds of mistake. One is doing work a platform has said in writing it does not need. The other is quieter and more expensive: assuming a technique keeps working after everybody adopts it.

Tian, Chen, Tang, Liu and Jia looked at where citations fail and repaired them by changing about five percent of a page. They also warn that optimising every page the same way costs you on the long tail.

What the advice says to doWhat Google publishesStatus
Publish an llms.txt or another AI text fileYou do not need to create new machine readable files, AI text files, markup, or Markdownstated unnecessary
Chunk content into small pieces for retrievalThere is no requirement to break your content into tiny piecesstated unnecessary
Add schema markup so AI can parse the pageStructured data is not required for generative AI searchstated unnecessary for this purpose
Write in a special register for machinesYou do not need to write in a specific way just for generative AI searchstated unnecessary
Quoted from Google Search Central, AI features and your website, updated 10 July 2026 and read 3 September 2026. Structured data still earns rich results in classic search; the claim being refused here is the AI one, although the advice rarely separates the two.

Access comes before tactics.

The file itself has other arguments for and against it, and they are worth separating from this one. llms.txt sets out what it is, who has adopted it and what the adoption numbers show.

Why not everyone can win

The second mistake has a measurement behind it. C-SEO Bench found that as the number of adopters increases, the overall gains decrease, describing the problem as congested and zero-sum. Put that beside the 115.1 and 30.3 figures from the founding experiment and the pattern holds. Within one answer, share moves between sources. Nothing is created.

And there is a step before any of this that most advice skips. On Cloudflare’s network a meaningful share of what AI bots ask for never reaches the page. Much of the crawling that does reach it is declared as training that never searches. Read the two distributions separately; adding them up hides the point. Access is not the last item on the checklist. For a store, the same access question shows up first in ordinary crawling, which ecommerce SEO covers, and the AI answer surfaces themselves are in AI Overviews.

Read from the live Cloudflare Radar page on 3 September 2026. The window rolls, so these shares move.
What AI crawlers ask for, and what comes back

The gap between platforms is the part that does not fit the usual framing. An answer engine reads hundreds of pages for every visitor it sends back, whereas a search engine reads a handful, and the two sit on the same axis below.

Read from the live Cloudflare Radar AI Insights page on 3 September 2026. Values were recorded from the published chart on 3 September 2026.
Pages crawled for every visitor referred back

What generative engine optimization tools sell you

Crawler access you can check yourself. Whether any of this moved your citations is the part that gets bought, and there is no way to compare two quotes for it. Four vendors publish four different units, and one publishes no price at all, so every side by side pricing table is comparing things that are not the same thing.

One product publishes both parts needed for a direct unit-cost calculation. Its plans list a price and the number of prompts each one includes, so you can divide one by the other before you buy.

Four vendor pricing pages, read 3 September 2026The unit each vendor sells, as published, set side by side

Quoted from the four published pricing pages on 3 September 2026. This is a comparison of the published units, not a picture of the pages.

The unit each vendor sells, as published, set side by side

You are buying a sample size.

Read the figure by its right-hand column, not its prices. One vendor sells prompts, one sells prompts bundled with engines, one sells credits that convert into answers at a rate it publishes, and one names four tiers and shows no figure at all. A price per prompt can be computed for the first, although for the last two it is undefined.

Why the unit decides the answer

The reason the unit matters is that the thing being bought is a sample of a system that does not repeat itself. Ronald Sielinski sampled three engines repeatedly: daily over nine days, then again at ten minute intervals. The same prompt returned different sources from one run to the next, often by a wide margin. His conclusion is that many apparent differences between sites fall inside that natural variation, so a number taken from a single run looks far more precise than it is.

So the question a store should ask is not which tool is cheapest per prompt. It is how many observations the plan buys, how often they repeat, and whether that is enough to tell a real movement from the noise the measurement itself produces. Where this work sits inside a wider programme is on the page for our answer engine optimization service.

How two studies count citations

Once the tools are running, the numbers they report need a definition. There is more than one, and the difference is not a detail. A page can be selected as a source and contribute nothing to the answer, and a brand can be mentioned constantly without any of it being attributable.

Zhang, He and Yao, arXiv 2604.25707, read 3 September 2026Two stages that get reported as one number

From the paper’s own description of its dataset and findings.

Two stages that get reported as one number

One word, two different events.

Their finding is that breadth and depth diverge. Perplexity and Google cite more sources on average, while ChatGPT cites fewer but shows substantially higher average influence among the pages it does fetch. The pages with high influence are longer, more structured, and richer in extractable evidence.

Semrush counts the same surface, yet arrives somewhere that looks opposite. Across 126 million United States prompts collected between January and April 2026, it reports ChatGPT citing an average of 15 sources per response against Gemini’s three. Both count cited sources and both include Gemini, so the two orderings disagree on the same pair. Neither study publishes enough to settle which is right. Different prompt sets, different months and different counting rules could each produce the gap, and none of them is documented in enough detail to test. The practical reading is narrower than either headline: a sources-per-response figure describes the measurement that produced it, not the engine. Any dashboard reporting one owes you the prompt set, the window and the counting rule, and the same question applies to the Google surfaces measured in AI search engine optimization.

Brand size as a fixed variable

The third variable is not something the tactics in this literature address at all. Pratyush Kumar analysed more than 100,000 prompt responses across over 100 brands between March and May 2026. How often a brand appears tracks how established it already is, in steps of roughly the same size.

From the paper’s reported first-run figures, read 3 September 2026.
How often a brand appears, by how well known it already is

For a small brand that is the size of the task. The tactics in this literature move share inside an answer you are already part of. Getting into the answer at all is the older problem, and the older tools still solve most of it. For a store that means the ordinary work in ecommerce SEO.

Open the page you most want quoted and look for the four things the absorbed pages are richer in: a definition, a number, a comparison, a short procedure. Nobody has tested whether adding the missing ones moves anything. Treat it as a bet on that association, not on a demonstrated cause. The bet is still smaller than any published checklist asks for.

Whether the number on your dashboard can even register that bet is a separate question, and the answer sits in the arithmetic: the score is a position-weighted mention rate over somebody else’s prompt set. What AI visibility can and cannot support is taken apart there.

Sources

  1. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande GEO: Generative Engine Optimization, KDD 24 proceedings, pages 5 to 16 24 August 2024
  2. Martinez Optimizing Visibility in Generative Engines: A Critical Survey of GEO 2023 to 2026, arXiv 2607.14035, preprint 15 July 2026
  3. Puerto, Gubri, Green, Oh and Yun C-SEO Bench: Does Conversational SEO Work? NeurIPS Datasets and Benchmarks 2025 6 June 2025
  4. Google Search Central AI features and your website updated 10 July 2026, accessed 3 September 2026
  5. Zhang, He and Yao From Citation Selection to Citation Absorption, arXiv 2604.25707, preprint 28 April 2026
  6. Kumar Generative Engine Optimization at Scale, arXiv 2606.20065, preprint 18 June 2026
  7. Sielinski Quantifying Uncertainty in AI Visibility, arXiv 2603.08924, preprint 9 March 2026
  8. Bagga, Farias, Korkotashvili, Peng and Wu E-GEO: A Testbed for Generative Engine Optimization in E-Commerce, arXiv 2511.20867, preprint 25 November 2025
  9. Tian, Chen, Tang, Liu and Jia Diagnosing and Repairing Citation Failures in Generative Engine Optimization, arXiv 2603.09296, preprint 10 March 2026
  10. Semrush 2026 AI Visibility Index, 126 million United States prompts 26 June 2026
  11. Cloudflare Radar AI Insights, crawl-to-refer ratio and crawl purpose, rolling seven day window read 3 September 2026
  12. Youell, Winsome Marketing GEO: What the KDD Paper Actually Proves and What It Does Not 3 March 2025 Measured 18 months ago, on a surface that has moved since.
  13. Otterly.AI Pricing read 3 September 2026
  14. Profound Pricing read 3 September 2026
  15. Rankscale Pricing read 3 September 2026
  16. Peec AI Pricing read 3 September 2026

Questions people ask

Is generative engine optimization a real thing?

GEO is a named method with a published measurement behind it. A 2024 KDD paper introduced the term, tested nine edits against a 10,000 query benchmark and reported that the best of them raised a document’s share of an answer from 19.3 to 27.2.

What is contested is how far that generalises. A 2025 benchmark tested the methods across six domains. Three of fifty-four combinations came out significantly positive, and none in question answering. Martinez’s 2026 survey of forty-five studies concluded that no reviewed technique shows a stable, cross-platform causal effect on discoverability.

Is GEO replacing SEO?

The evidence points the other way. C-SEO Bench reports that traditional SEO strategies, the ones that improve where a source ranks in the model’s context, are significantly more effective than the GEO methods it tested.

The guidance Google publishes is that the best practices for SEO continue to be relevant, because these AI features rely on core Search ranking systems. A page has to be indexed and eligible to appear with a snippet before anything else applies.

What is the best tool for generative engine optimization?

That question has no answerable version from published prices, because the products do not sell the same unit. Published plans use incompatible units: a prompt allowance, prompts tied to one engine, or credits convertible into answers. One vendor publishes no price at all.

The useful comparison is how many observations a plan buys and how often they repeat, because repeated sampling of the same prompt produces different citations.

What’s the difference between SEO and GEO?

The difference is which stage is being measured. SEO work aims at whether a page is retrieved and where it ranks. GEO work, as the founding paper defined it, aims at how much of a generated answer a document contributes once it is already in the context.

That is why an edit can help one and hurt the other. Different teams measured the gains in the first setting and the ranking harm in the second. The two sit next to each other; neither is inside the other.