Generative engine optimization (GEO) means editing web content so that AI systems which write answers, such as ChatGPT, Perplexity and Google's AI Overviews, cite it as a source. (It has nothing to do with geographic or geo-targeted SEO.) The term comes from a 2023 research paper that reported visibility gains of up to 40%, and that number has been quoted in nearly every GEO pitch since.
So does it work? The best evidence so far says: mostly not, at least not the specific content tactics the paper tested. A larger 2025 benchmark, C-SEO Bench, retested eight of those methods on four language models and found most of them ineffective and some harmful. What moved citations far more than any rewrite was simply where a page sat in the list of documents the model was given.
Below is what each study did, the numbers side by side, and what's still worth doing.
The 2023 paper: where "up to 40%" comes from
"GEO: Generative Engine Optimization" by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande (Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi) was posted to arXiv in November 2023 and published at KDD 2024.
The authors built GEO-bench, 10,000 queries each paired with the sources an engine would draw on, and tested nine content modifications against a generative engine designed to mimic Bing Chat. The abstract's claim: "GEO can boost visibility by up to 40% in generative engine responses." When the strongest methods were checked on Perplexity, gains reached up to 37%. The three best were Cite Sources, Quotation Addition and Statistics Addition. Keyword stuffing was among the weakest.
Two features of the setup matter for what came next. Only one source was being optimized at a time, and the engine was a single 2023-era architecture.
The 2025 replication: C-SEO Bench
"C-SEO Bench: Does Conversational SEO Work?" by Haritz Puerto, Martin Gubri, Tommaso Green, Seong Joon Oh and Sangdoo Yun (Parameter Lab, NAVER AI Lab, TU Darmstadt's UKP Lab, the University of Mannheim and the University of Tübingen) was accepted at the NeurIPS 2025 Datasets and Benchmarks track.
It set out to test conversational search optimization systematically:
- Ten methods. Eight from the GEO paper (Authoritative, Statistics, Citations, Fluency, Unique Words, Technical Terms, Simple Language, Quotes), which is every GEO method except keyword stuffing, plus two new ones (Content Improvement and LLM Guidance).
- 1,921 queries in six domains and two tasks. Product recommendation (retail, video games, books) and question answering (web, news, debate), over 16,360 documents.
- Four models. GPT-4o-mini as the main one, plus Claude 3.5 Haiku, o3 and o4-mini.
- Several optimizers at once. Besides the single-actor setting, it varied how many competing documents used the same method.
The headline from the abstract: "most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking."
The numbers side by side
C-SEO Bench compared each content method with a plain retrieval baseline: put the target document first in the list the model reads. Here is the retail task, the clearest case, for the two models the paper reports in Tables 3 and 4. Values are the average improvement in the document's citation rank (higher is better), with the paper's standard deviations.
| Model (table in the paper) | Document placed first | Best content method | Ratio |
|---|---|---|---|
| GPT-4o-mini (Table 3) | +2.77 (±2.31) | +0.36 (±1.47), LLM Guidance | ≈7.7× |
| Claude 3.5 Haiku (Table 4) | +1.61 (±1.96) | −0.29 (±1.28), Content Improvement | best method hurt |
Two things stand out. The "7.7×" figure you'll see quoted holds for one model in one task. And on Claude 3.5 Haiku the best-performing content method still made things slightly worse on average. The spreads are wide, so individual pages can move either way, but the averages don't support the idea that rewriting for GEO reliably earns citations.
The authors' own summary: "making the target document the first one in the LLM context leads to far greater citation ranking gains in the LLM response than any C-SEO method."
The advantage shrinks as others copy you
C-SEO Bench also varied how many competing documents applied the same method. Its finding (section 6.4): "while early adopters experience significant gains, these gains decrease steadily as more players adopt the same method." The authors describe the setting as "congested and zero-sum."
This is the finding with the most practical weight. A formatting trick that works because few people use it stops working once a tool applies it for everyone. Changes that make a page more useful to a reader don't decay the same way, because they were never relative to competitors in the first place.
Why the two studies disagree
My reading, not either paper's claim: they measured different situations. The GEO paper asked what happens when one page in a set is improved and nothing else changes. C-SEO Bench asked what happens across several newer models, in more domains, and when competitors use the same playbook. The second setup is closer to the web in 2026, where any content tool can apply "add statistics, add quotes" at scale.
Neither study measured a live consumer product. Both gave models a fixed set of documents, which is a reasonable lab stand-in for the synthesis step but says nothing about how ChatGPT or Google pick those documents in the first place. That earlier step is retrieval, and it's where C-SEO Bench found the big lever.
What's still worth doing
Sorted by how much evidence stands behind each.
Get retrieved, near the top. The one lever C-SEO Bench measured as strong. In practice that means ordinary SEO: indexed pages, crawlers allowed, and rankings for the query and its sub-queries. Engines search different indexes (Google's for AI Overviews; for ChatGPT, a blend of OpenAI's own index and third-party search results), so check each engine's crawlers: OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot and Bingbot, plus the user-triggered fetchers covered in what is AEO.
Cite real sources, for readers. Attribution was the GEO paper's strongest family of methods; in C-SEO Bench, no content method, this one included, came close to the effect of retrieval position. Do it anyway: a linked source lets a reader check you, costs nothing, and doesn't depend on any study being right.
Write clearly. Fluency and simple language were among the GEO methods, with the same caveat. Clear writing is still easier to quote and to trust, and it has never stopped being good SEO.
Skip keyword stuffing. It was among the worst performers in the original benchmark. It hasn't been retested since (C-SEO Bench left it out), so this rests on one study plus long SEO experience.
Not tested by either paper: answer-first openings, schema markup and named authorship. These are practitioner consensus with plausible mechanisms, covered in what is AEO. Treat articles calling them "validated by the GEO study" as stretching the citation.
Run the test on your own pages
Given how differently the two papers came out, the most useful evidence is your own. A small, honest test is cheap to set up:
- Pick 20 prompts your audience actually asks, and the pages you'd want cited for them.
- Record a baseline: run each prompt several times in each engine you care about for two weeks, and log whether your domain is cited.
- Change five pages (for example, add sourced statistics and an answer-first opening) and leave five comparable pages alone as controls.
- Keep sampling on the same schedule for six weeks or more, then compare the change in citation rate between the two groups.
Ten pages won't give you a publishable effect size, but it will tell you whether a tactic is worth rolling out across the site, on the engines and topics that matter to you.
FAQ
Is GEO the same as AEO?
Did C-SEO Bench test ChatGPT or Google directly?
Should I stop doing GEO?
Sources
- Aggarwal et al. — GEO: Generative Engine Optimization, KDD 2024 (arXiv 2311.09735, November 2023)
- Puerto et al. — C-SEO Bench: Does Conversational SEO Work?, NeurIPS 2025 Datasets and Benchmarks (arXiv 2506.11097); tables and section 6.4 in the full HTML version (v3)
- Peec AI — ChatGPT built its own search index, September 2026
Updated October 1, 2026: refocused on the replication evidence; corrected the C-SEO Bench method count (ten, eight from GEO) and ratio (about 7.7×, GPT-4o-mini only) and added per-model results.