Does GEO Actually Work? What the Peer-Reviewed Research Says
26 August 2026 · GEONI Team
“You can increase your visibility in AI answers” is said constantly, usually with nothing behind it. But the claim does have a peer-reviewed source: GEO: Generative Engine Optimization, published at ACM SIGKDD in 2024. We put the same question to live engines in 2026, across 154 subjects. This page sets the two side by side: what the research found, what we measured, and where they meet. Research figures come straight from the paper; GEONI figures from our own measurement.
What the study did
The researchers built a benchmark called GEO-bench: 10,000 queries, drawn from nine different sources, spread across 25 domains. They put those queries to a generative engine, recorded which sources the answers cited, then modified the content of a single source and asked again. What is measured is whether that modification raised or lowered that source's visibility.
The design matters: this is a controlled experiment. Unlike “we did this and our traffic went up” case studies, one variable changes and its effect is isolated.
How visibility was measured
Two metrics were defined, because merely appearing in an answer means little on its own:
- Position-Adjusted Word Count — the word count of sentences attributable to a source, weighted by where in the answer they appear. Being discussed at length up front is not the same as a single mention at the bottom.
- Subjective Impression — a language model's judgement of how much the source dominates the answer, across dimensions such as relevance, uniqueness and influence.
Results: what worked and what did not
Nine methods were tested. Untouched content scored a baseline of 19.3. The figures below are the paper's Overall column:
| Method | Score | Verdict |
|---|---|---|
| Quotation Addition | 27.2 | Highest |
| Statistics Addition | 25.2 | Strong |
| Fluency Optimization | 24.7 | Strong |
| Technical Terms | 22.7 | Positive |
| Easy-to-Understand | 22.0 | Positive |
| Authoritative | 21.3 | Positive |
| Unique Words | 20.5 | No real effect |
| Keyword Stuffing | 17.7 | BELOW baseline |
In the paper's own words, the three top methods — Cite Sources, Quotation Addition and Statistics Addition — achieved a relative improvement of 30–40% on Position-Adjusted Word Count and 15–30% on Subjective Impression. The best methods beat the baseline by 41% and 28% respectively.
The striking result: an old SEO tactic actively hurt
Keyword stuffing performed worse than doing nothing at all (17.7 against 19.3). The study files it directly under non-performing methods.
The practical reading is clean: density-based tactics that once worked in ranking-era search find no purchase in a system that generates the answer. A generative engine does not rank your page, it quotes your sentence — and a stuffed sentence is not a quotable one.
GEONI measurement · July–August 2026
The study says “be quotable”. We measured that most sites have not built the structure to be quoted from.
The paper shows the winning methods share one property: verifiability. So where do sites actually stand? Across the 100 websites we measured:
Zero of one hundred blocked AI crawlers. Access is not the constraint. The gap is not a locked door, it is an empty room: the structure an engine could parse and quote was never built.
The real caveat: there is no single recipe
This is the study's most frequently skipped finding. The effect of the same method varies dramatically by domain. Quotation Addition, across the domains measured, reduced visibility by 22.9% at one end and increased it by 99.7% at the other.
That is why the authors explicitly argue for domain-specific optimisation. Any listicle telling you “do these five things to appear in AI” is stepping over this result. The defensible approach is not to apply the general recipe but to measure in your own domain.
GEONI measurement · 90 subjects, four engines
We measured the same swing between engines — and the distribution came out bimodal.
For each subject we took the gap between its highest- and lowest-scoring engine. We expected a spread around some average:
- 0–19 points (engines broadly agree) — 40 subjects
- 20–39 points (partial disagreement) — 4 subjects
- 40+ points (fundamental disagreement) — 46 subjects
Only 4 of 90 sit in the middle. A brand is either seen roughly the same way everywhere or completely differently — there is no gradual middle. Median gap 46.8; largest measured 91.5: same subject, same day, near-complete in one engine and near-zero in another.
The study says the effect varies by domain. Our measurement adds a second axis: it varies by engine too, and without gradation.
Limits of the study — and what we add
Reading a source honestly means stating where it stops:
- Controlled setting. Measurement ran on the researchers' generative-engine setup. It is not identical to how ChatGPT or Gemini behave live today.
- Timing. The work was prepared in late 2023 and published in 2024. Models have been updated many times since.
- Relative metrics. The results describe share of visibility inside an answer, not customers acquired. Commercial outcome is a separate question this study does not measure.
- One source optimised. In reality competitors optimise simultaneously; the experiment does not model that.
GEONI measurement · repeat scans
Even when no competitor moves, the number does not stay still.
We measured the same subject four times on the same day, within 83 minutes. Whether the engines recognised it flipped between runs; the score went 34 → 17 → 15 → 65. Nothing had changed — not the content, not the competitors.
The practical consequence is blunt: one reading is not a measurement. What makes a number trustworthy is repeat sampling, a dated log of every run, and smoothing across runs. A tool that shows you one number with no history is showing you a coin flip.
What it means in practice
Both sources point the same way. The research says what to write: support the claim with a number, show where it came from, ground it in a credible quotation, write the sentence clearly. Our measurement says where people actually stand: most sites have not built even the structure that precedes all of that, and looking at one engine — or one reading — misleads.
But the study's own warning matters most: the effect depends on the domain. Nothing on this page substitutes for measuring in yours.
Source
Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A. (2024). GEO: Generative Engine Optimization. Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '24), Barcelona.
- Paper (arXiv): arxiv.org/abs/2311.09735
- Proceedings (ACM): dl.acm.org/doi/10.1145/3637528.3671900
GEONI measurement: 9 July – 24 August 2026, 154 subjects (90 with a complete four-engine measurement). Only aggregates are published; no customer name, domain or individual score is disclosed. The sample skews Turkish-language and is not a global average.
All figures on this page are taken directly from the paper above, not relayed from secondary summaries. If you spot an error, tell us: mail@geoni.ai
Measure in your own domain
The study's conclusion in one line: there is no general recipe, only domain-specific measurement. GEONI puts your category questions to ChatGPT, Claude, Gemini and Perplexity separately, extracts the sources those answers cite, and audits whether your site is machine-readable. Your first scan is free.
Run a free scan →