Where generative engine optimization came from
The definition of GEO is not in dispute. Where the word came from matters more. The paper that introduced it measured something specific, and it measured it in a narrower situation than the number sounds.
The word has a paper behind it.
Generative engine optimization was named in a paper by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande. It was published in the proceedings of the 30th ACM SIGKDD conference on 24 August 2024. The paper is open access and runs from page 5 to page 16.
What that number measures and what it does not
The paper introduced GEO-bench, a benchmark of 10,000 queries. It measured visibility with position-adjusted word count: how much of a generated answer traces back to a given source, weighted so earlier sentences count for more. That is a share-of-answer measure. It is not traffic, not ranking, and not whether the engine found you.
The distinction matters because of how the experiment was built. Five documents were already placed in the model’s context, and the question was which of them the answer would draw from. The consequence is straightforward. A forty percent gain there does not mean forty percent more traffic, forty percent higher ranking or forty percent more clicks. It means the optimised source contributed more to the answer in that pipeline.
Everything about how Google’s own AI surfaces behave, and how far citation has drifted from ranking on them, is a separate question with its own measurements. AI search engine optimization takes those apart. What comes next is the method itself and the testing it has produced since. The shorter definition of its sibling term sits in answer engine optimization, and what the GEO tools sold under this name buy you is a separate page.
How nine tested edits performed
The method the paper proposed is not a philosophy. It is nine concrete edits to a document, each one measured separately, which is why the testing that followed could argue with it. The results are more interesting than the headline, because they point in a consistent direction.
Read the order, not the size.
The survey that reviewed this literature draws two readings out of that ordering, and both are its own. First, directly extractable information, meaning figures, definitions, quotations and references, can facilitate the use of a document. Second, in its words, keyword stuffing does not transfer effectively from conventional SEO. What the ranking on its own establishes is the order, whereas the reason for it stays an interpretation. The order is the useful part. The cheap edits are the ones that add something a reader could check, not the ones that reword what is already there.
A second result rarely quoted
There is a second result inside the same experiment that rarely gets quoted, and it changes what the advice means. When Cite Sources was applied, the fifth source in the context gained 115.1 percent while the first lost 30.3 percent. The metric is a share of one answer, so it is competitive by construction. Whatever the fifth source gained, it gained from somebody.
A short definition of the neighbouring term, and where the two words came from, sits in answer engine optimization.
What a benchmark found across six domains
Two years after the paper, the obvious question had an answer. Haritz Puerto, Martin Gubri, Tommaso Green, Seong Joon Oh and Sangdoo Yun built C-SEO Bench and published it at NeurIPS Datasets and Benchmarks in 2025. It tests the methods across two tasks and three domains each, including product recommendation.
The three findings this page leans on, in one paragraph
- 1Ineffective is the weaker of the two words. The sentence says frequently negative, so a method can move a page down rather than leave it where it was.
- 2Reviewed, and the frame is where you can tell: accepted at NeurIPS Datasets and Benchmarks 2025. Most of the other work cited on this page is a preprint.
- 3Two search tasks and three domains each, so the six domains are the breadth of the test and also its limit.
On stores the result is the same shape. E-GEO, by Bagga, Farias, Korkotashvili, Peng and Wu, ran 13,747 product queries against five generative engines.
Martinez’s 2026 critical survey reports that ten of E-GEO’s fifteen starting heuristics came out neutral or negative. The count is the survey’s, not the testbed authors’.
Two years later, the answer is narrow.
Most current methods are not only largely ineffective but frequently have a negative impact on document ranking, and that is the paper’s own summary of its results. Several transformations reduce rank outright. What matters is the shape of the test behind that sentence: how many methods, in how many domains, and which task produced nothing at all.
The founding number and the benchmark that tested it belong side by side, and the two are not in conflict once you notice they measure different stages. The paper measures how much of an answer a document contributes once it is already in the context. The benchmark measures whether the document gets into the ranking at all. An edit could raise the first and lower the second. That is the shape the two literatures suggest together, though no single experiment measured both on the same edit.
Where the evidence points instead
That leaves a narrow, boring instruction, and it comes from the research. The same benchmark reports that traditional SEO strategies, the ones aiming to improve the ranking of the source in the model’s context, are significantly more effective. Google puts it from the other side. Its own guidance says the best practices for SEO continue to be relevant, because these features rely on core Search ranking systems. The e-commerce testbed lands there too. Across 13,747 product queries, the survey reviewing it reports that ten of the fifteen hand-written heuristics came out neutral or negative. What beat them was optimisation done systematically, not from a list, which is the same conclusion the answer engine optimization entry reaches about checklists.
-
Get retrievable first
Google states the eligibility requirement plainly: a page has to be indexed and eligible to be shown with a snippet. Nothing downstream matters until that holds, and on a store the usual blockers are the ones an ecommerce SEO audit looks for.
-
Put extractable material on the page
Zhang and colleagues found that pages with high citation influence are longer, more structured and richer in definitions, numerical facts, comparisons and procedural steps. That is the same direction the paper’s three winning edits point in.
-
Change little, and change the right thing
The AgentGEO work reports over 40 percent relative improvement in citation rates while modifying only 5 percent of a document, against 25 percent for baseline methods. The same paper warns that generic optimisation can harm long-tail content.
-
Stop before the checklist
The standard GEO list is where the evidence runs out, and some items on it have been called unnecessary by the platforms themselves.
If you want crawler access, indexing and structure checked instead of argued about, an ecommerce SEO audit covers exactly that.
What the platforms call unneeded
Crawler access, indexing and structure are the parts you can have checked. Past them sit two kinds of mistake. One is doing work a platform has said in writing it does not need. The other is quieter and more expensive: assuming a technique keeps working after everybody adopts it.
Tian, Chen, Tang, Liu and Jia looked at where citations fail and repaired them by changing about five percent of a page. They also warn that optimising every page the same way costs you on the long tail.
| What the advice says to do | What Google publishes | Status |
|---|---|---|
| Publish an llms.txt or another AI text file | You do not need to create new machine readable files, AI text files, markup, or Markdown | stated unnecessary |
| Chunk content into small pieces for retrieval | There is no requirement to break your content into tiny pieces | stated unnecessary |
| Add schema markup so AI can parse the page | Structured data is not required for generative AI search | stated unnecessary for this purpose |
| Write in a special register for machines | You do not need to write in a specific way just for generative AI search | stated unnecessary |
Access comes before tactics.
The file itself has other arguments for and against it, and they are worth separating from this one. llms.txt sets out what it is, who has adopted it and what the adoption numbers show.
Why not everyone can win
The second mistake has a measurement behind it. C-SEO Bench found that as the number of adopters increases, the overall gains decrease, describing the problem as congested and zero-sum. Put that beside the 115.1 and 30.3 figures from the founding experiment and the pattern holds. Within one answer, share moves between sources. Nothing is created.
And there is a step before any of this that most advice skips. On Cloudflare’s network a meaningful share of what AI bots ask for never reaches the page. Much of the crawling that does reach it is declared as training that never searches. Read the two distributions separately; adding them up hides the point. Access is not the last item on the checklist. For a store, the same access question shows up first in ordinary crawling, which ecommerce SEO covers, and the AI answer surfaces themselves are in AI Overviews.
The gap between platforms is the part that does not fit the usual framing. An answer engine reads hundreds of pages for every visitor it sends back, whereas a search engine reads a handful, and the two sit on the same axis below.
What generative engine optimization tools sell you
Crawler access you can check yourself. Whether any of this moved your citations is the part that gets bought, and there is no way to compare two quotes for it. Four vendors publish four different units, and one publishes no price at all, so every side by side pricing table is comparing things that are not the same thing.
One product publishes both parts needed for a direct unit-cost calculation. Its plans list a price and the number of prompts each one includes, so you can divide one by the other before you buy.
Four vendor pricing pages, read 3 September 2026The unit each vendor sells, as published, set side by side
Quoted from the four published pricing pages on 3 September 2026. This is a comparison of the published units, not a picture of the pages.

You are buying a sample size.
Read the figure by its right-hand column, not its prices. One vendor sells prompts, one sells prompts bundled with engines, one sells credits that convert into answers at a rate it publishes, and one names four tiers and shows no figure at all. A price per prompt can be computed for the first, although for the last two it is undefined.
Why the unit decides the answer
The reason the unit matters is that the thing being bought is a sample of a system that does not repeat itself. Ronald Sielinski sampled three engines repeatedly: daily over nine days, then again at ten minute intervals. The same prompt returned different sources from one run to the next, often by a wide margin. His conclusion is that many apparent differences between sites fall inside that natural variation, so a number taken from a single run looks far more precise than it is.
So the question a store should ask is not which tool is cheapest per prompt. It is how many observations the plan buys, how often they repeat, and whether that is enough to tell a real movement from the noise the measurement itself produces. Where this work sits inside a wider programme is on the page for our answer engine optimization service.
How two studies count citations
Once the tools are running, the numbers they report need a definition. There is more than one, and the difference is not a detail. A page can be selected as a source and contribute nothing to the answer, and a brand can be mentioned constantly without any of it being attributable.
Zhang, He and Yao, arXiv 2604.25707, read 3 September 2026Two stages that get reported as one number
From the paper’s own description of its dataset and findings.

One word, two different events.
Their finding is that breadth and depth diverge. Perplexity and Google cite more sources on average, while ChatGPT cites fewer but shows substantially higher average influence among the pages it does fetch. The pages with high influence are longer, more structured, and richer in extractable evidence.
Semrush counts the same surface, yet arrives somewhere that looks opposite. Across 126 million United States prompts collected between January and April 2026, it reports ChatGPT citing an average of 15 sources per response against Gemini’s three. Both count cited sources and both include Gemini, so the two orderings disagree on the same pair. Neither study publishes enough to settle which is right. Different prompt sets, different months and different counting rules could each produce the gap, and none of them is documented in enough detail to test. The practical reading is narrower than either headline: a sources-per-response figure describes the measurement that produced it, not the engine. Any dashboard reporting one owes you the prompt set, the window and the counting rule, and the same question applies to the Google surfaces measured in AI search engine optimization.
Brand size as a fixed variable
The third variable is not something the tactics in this literature address at all. Pratyush Kumar analysed more than 100,000 prompt responses across over 100 brands between March and May 2026. How often a brand appears tracks how established it already is, in steps of roughly the same size.
For a small brand that is the size of the task. The tactics in this literature move share inside an answer you are already part of. Getting into the answer at all is the older problem, and the older tools still solve most of it. For a store that means the ordinary work in ecommerce SEO.
Open the page you most want quoted and look for the four things the absorbed pages are richer in: a definition, a number, a comparison, a short procedure. Nobody has tested whether adding the missing ones moves anything. Treat it as a bet on that association, not on a demonstrated cause. The bet is still smaller than any published checklist asks for.
Whether the number on your dashboard can even register that bet is a separate question, and the answer sits in the arithmetic: the score is a position-weighted mention rate over somebody else’s prompt set. What AI visibility can and cannot support is taken apart there.
Sources
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande GEO: Generative Engine Optimization, KDD 24 proceedings, pages 5 to 16
- Martinez Optimizing Visibility in Generative Engines: A Critical Survey of GEO 2023 to 2026, arXiv 2607.14035, preprint
- Puerto, Gubri, Green, Oh and Yun C-SEO Bench: Does Conversational SEO Work? NeurIPS Datasets and Benchmarks 2025
- Google Search Central AI features and your website
- Zhang, He and Yao From Citation Selection to Citation Absorption, arXiv 2604.25707, preprint
- Kumar Generative Engine Optimization at Scale, arXiv 2606.20065, preprint
- Sielinski Quantifying Uncertainty in AI Visibility, arXiv 2603.08924, preprint
- Bagga, Farias, Korkotashvili, Peng and Wu E-GEO: A Testbed for Generative Engine Optimization in E-Commerce, arXiv 2511.20867, preprint
- Tian, Chen, Tang, Liu and Jia Diagnosing and Repairing Citation Failures in Generative Engine Optimization, arXiv 2603.09296, preprint
- Semrush 2026 AI Visibility Index, 126 million United States prompts
- Cloudflare Radar AI Insights, crawl-to-refer ratio and crawl purpose, rolling seven day window
- Youell, Winsome Marketing GEO: What the KDD Paper Actually Proves and What It Does Not Measured 18 months ago, on a surface that has moved since.
- Otterly.AI Pricing
- Profound Pricing
- Rankscale Pricing
- Peec AI Pricing
Questions people ask
Is generative engine optimization a real thing?
GEO is a named method with a published measurement behind it. A 2024 KDD paper introduced the term, tested nine edits against a 10,000 query benchmark and reported that the best of them raised a document’s share of an answer from 19.3 to 27.2.
What is contested is how far that generalises. A 2025 benchmark tested the methods across six domains. Three of fifty-four combinations came out significantly positive, and none in question answering. Martinez’s 2026 survey of forty-five studies concluded that no reviewed technique shows a stable, cross-platform causal effect on discoverability.
Is GEO replacing SEO?
The evidence points the other way. C-SEO Bench reports that traditional SEO strategies, the ones that improve where a source ranks in the model’s context, are significantly more effective than the GEO methods it tested.
The guidance Google publishes is that the best practices for SEO continue to be relevant, because these AI features rely on core Search ranking systems. A page has to be indexed and eligible to appear with a snippet before anything else applies.
What is the best tool for generative engine optimization?
That question has no answerable version from published prices, because the products do not sell the same unit. Published plans use incompatible units: a prompt allowance, prompts tied to one engine, or credits convertible into answers. One vendor publishes no price at all.
The useful comparison is how many observations a plan buys and how often they repeat, because repeated sampling of the same prompt produces different citations.
What’s the difference between SEO and GEO?
The difference is which stage is being measured. SEO work aims at whether a page is retrieved and where it ranks. GEO work, as the founding paper defined it, aims at how much of a generated answer a document contributes once it is already in the context.
That is why an edit can help one and hurt the other. Different teams measured the gains in the first setting and the ranking harm in the second. The two sit next to each other; neither is inside the other.