What is GEO? Generative Engine Optimization, and what the research actually found
GEO is the practice of optimizing content for citation in AI answer engines. The term comes from a 2023 paper that measured which content changes increase citation — and reported gains of up to 40%. Later benchmarks failed to reproduce most of it. Here's what the original paper actually said, what has since been contradicted, and what still holds.
GEO (Generative Engine Optimization) is the practice of structuring web content so AI answer engines cite it — ChatGPT with search, Perplexity, Google AI Overviews, Gemini. The term comes from a paper published at KDD 2024 by researchers from Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi, which tested nine content modifications across a 10,000-query benchmark and reported visibility gains of up to 40% for the strongest three. That headline number is quoted constantly and almost never qualified. It should be: later benchmarks have failed to reproduce most of these tactics, and at least one found them performing below an unmodified baseline. GEO overlaps almost entirely with AEO — the terms are interchangeable in practice.
Most articles about GEO stop at the 40% figure. This one goes further, because the more useful question in 2026 isn't what the original paper claimed — it's what has held up since.
Where the term came from
"Generative Engine Optimization" was introduced in a paper titled GEO: Generative Engine Optimization (Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan et al.), posted to arXiv in November 2023 and published at KDD 2024. It was the first formal academic treatment of optimizing content for engines that synthesize an answer rather than return a list of links.
The researchers built GEO-bench — a benchmark of 10,000 queries across multiple domains, each paired with the web sources an engine would draw on. They tested nine content modification strategies using a generative engine built to mimic Bing Chat, then validated the strongest tactics on Perplexity as a real-world check.
The paper introduced two metrics, and the distinction matters for reading any summary of it:
- Position-Adjusted Word Count — how much of the answer comes from your source, weighted by where the citation appears
- Subjective Impression — a composite score of how the source is represented
Almost every secondhand summary quotes gains against the first metric while omitting the second, which is consistently lower.
What the paper actually measured
The headline: the three strongest methods — Cite Sources, Quotation Addition and Statistics Addition — achieved a relative improvement of 30–40% on Position-Adjusted Word Count and 15–30% on Subjective Impression.
That is a range for three methods across two metrics. It is not a per-tactic figure, and any article claiming "statistics improved citation by 37%" or "quotes improved it by 41%" is quoting a number that does not appear in the paper.
The most interesting result is the last one. Keyword stuffing performed 8% below the unmodified baseline on Position-Adjusted Word Count, and 10% below in the Perplexity validation. Loading a page with query terms — the oldest reflex in search optimization — measurably reduced how often a generative engine cited it.
Fluency optimization is worth noting too: roughly a 28% gain from editing quality alone, with no new information added. Clear prose appears easier for a model to parse, summarize and attribute.
What later research found
This is the part missing from most GEO articles, and the reason to be careful with the 40% number.
C-SEO Bench (Puerto et al., 2025) — the first systematic benchmark of conversational-search optimization tactics — concluded that most of these tactics do not help, several actively hurt, and plain source relevance keeps working. That is close to the opposite of how the original findings are usually marketed.
A more recent evaluation went further: across GEO-Bench's scale and diversity, token-level GEO heuristics failed to consistently improve citation visibility over unmodified text, and in several cases scored below it. On one engine, baseline visibility was 13.34% while the heuristic variants landed between 10.92% and 12.21%. Some also degraded content quality scores.
This is not a reason to ignore GEO. It is a reason to distrust anyone selling it with a fixed number attached. The mechanisms below still hold; the multipliers do not.
How generative citation actually works
Three stages sit between your page and a citation. Understanding which stage you are failing at matters more than any single tactic.
Stage one is the part GEO articles tend to wave away. A page that is not indexed, not crawlable by that engine's bot, or nowhere near the top of the underlying results will not be cited regardless of how well-structured it is. GEO is a layer on SEO, not an alternative to it.
Stage three is why extractability matters. Pages built from discrete units — FAQ pairs, comparison tables, definition blocks — offer cleaner things to quote than the same information delivered as continuous prose.
What still holds, and on what evidence
Separating the two is the whole point of this article.
Measured in the original paper, contested since:
- Cite your sources. Attribution to named studies, tools and companies. The strongest single family of tactics in the paper.
- Add statistics. Specific numbers rather than vague quantities. With the caveat this article exists to make: only numbers you can attribute.
- Quote authorities. Exact quotes with attribution, not paraphrase.
- Write well. ~28% from editing alone, and the least likely finding to be architecture-specific.
- Do not keyword stuff. Measurably negative. The one finding no later work has contradicted.
Not in the paper — practitioner consensus, no controlled measurement:
- Direct answers in the opening. Widely reported, mechanically plausible given stage three, unmeasured at benchmark scale.
- Schema markup. FAQPage, Article with a real Person author, Review for commerce. Necessary rather than sufficient.
- Named authorship. Anonymous content appears to be filtered harder, especially in YMYL topics.
The second group may well be right. It is not research, and articles presenting it as "validated by the GEO study" — as an earlier version of this page did — are extending a citation past what it covers.
GEO and AEO are the same thing
For execution purposes, yes. The checklist is identical: direct-answer formatting, schema markup, named author signals, structured content, factual specificity, recency.
The difference is emphasis and origin. GEO stresses that these engines generate text, and carries academic weight from the paper. AEO stresses that they answer questions, and predates the paper by a year or two, coming out of voice search and featured snippet work. For a fuller treatment of the practice itself, see our complete AEO guide.
One clarification worth making: some sites use "GEO" to mean geographic or local SEO. In AI search discussions in 2026, it almost always means Generative Engine Optimization.
What to do with this
If you are starting today, the work is the same whichever term you use:
- Fix retrieval first. Confirm your pages are indexed and that the engines' crawlers can reach them. Check robots.txt for OAI-SearchBot, PerplexityBot, Claude-SearchBot and Bingbot — blocking a retrieval crawler removes you from that engine's answers entirely.
- Rewrite openings. The first 50–80 words should answer the page's question completely, with no preamble.
- Add schema. Article with a real Person author, FAQPage on question-style pages, Organization at site level.
- Attribute everything. Every number gets a named, linkable source. This is the paper's strongest finding and — unlike the multipliers — it costs nothing if it turns out to be weaker than reported.
- Stop keyword stuffing. The one thing measured to actively hurt.
- Measure your own citation rate. Ten to twenty representative queries, run manually across engines, weekly. Given how poorly the published effects have replicated, your own data on your own pages is worth more than anyone's benchmark.
FAQ
Is GEO the same as AEO?
What did the original GEO paper actually measure?
Are the GEO paper's findings still reliable in 2026?
Does GEO replace SEO?
What's the one GEO finding nobody has contradicted?
Closing
GEO is a real discipline with a weaker evidence base than its marketing suggests. One paper measured real effects on one engine architecture in 2023; the field has quoted its headline number ever since, largely without checking whether it replicated. It mostly has not.
That does not make the work pointless — it makes the framing matter. Cite your sources because attribution is verifiable and useful, not because a study promised 40%. Write clearly because clear writing is easier to quote. Structure content into extractable units because that is mechanically how stage three works. Skip keyword stuffing because it is the one tactic measured to backfire.
And measure your own pages. Given how unevenly published findings have held up, ten queries you run yourself every week are worth more than any benchmark someone else ran two years ago on a different engine.
Sources cited in this piece
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan et al., "GEO: Generative Engine Optimization" — arXiv, November 2023; published at KDD 2024. Source of the 30–40% / 15–30% figures, the nine tested tactics, GEO-bench, and the keyword-stuffing result. arxiv.org/pdf/2311.09735
- Puerto et al., "C-SEO Bench," 2025 — systematic benchmark of conversational-search optimization tactics; finding that most do not help and several hurt.
- Subsequent GEO-Bench evaluation — finding that token-level GEO heuristics fail to consistently improve citation visibility over an unmodified baseline (13.34% baseline vs 10.92–12.21% for heuristic variants on one engine).
Last updated 31 July 2026: corrected the reported effect sizes against the primary source, removed per-tactic figures that do not appear in the paper, and added the replication section.