How Four AI Engines See the Same Brand
25 August 2026 · GEONI Team · Measurement window: 9 July – 24 August 2026
Aggregated results from 154 unique measured subjects — websites, brands, people and social accounts — scanned across ChatGPT, Claude, Gemini and Perplexity. Every number on this page comes from our own measurement. There are no industry estimates and no figures quoted from anyone else. No customer name, domain or individual score appears anywhere.
Finding 1 — the most used engine recognises the fewest brands
| Engine | Recognised the subject | Hallucination suspicion | Judge accuracy score | Average visibility |
|---|---|---|---|---|
| Perplexity | 89% | 19% | 75.2 | 65.6 |
| Claude | 84% | 13% | 76.0 | 59.7 |
| Gemini | 90% | 29% | 68.6 | 59.2 |
| ChatGPT | 67% | 12% | 72.9 | 45.7 |
The engine with the largest user base recognised the fewest subjects. ChatGPT knew 67% of what we measured; Gemini knew 90%. The practical consequence: being known to Gemini tells you nothing about being known to ChatGPT. A brand that checks its visibility in one chat window and concludes it is fine has measured one quarter of the surface.
Finding 2 — the engine that recognises most also invents most
Every answer was compared against live web data by an independent judge model looking for contradictions and unverifiable claims. Gemini carried hallucination suspicion in 29% of measurements — roughly 2.4× ChatGPT's 12%. Gemini also had the lowest judge accuracy score at 68.6, while Claude had the highest at 76.0.
The conclusion runs opposite to intuition: being recognised and being described correctly are different properties. An engine mentioning your brand is not evidence that it is saying true things about you, and the engine most likely to mention you is the one most likely to be wrong.
Finding 3 — engines either agree closely or disagree wildly, with almost nothing in between
For each subject measured on all four engines we took the gap between the highest and lowest engine score. We expected a normal-looking spread around some average. What we found was bimodal:
| Gap between best and worst engine | Subjects |
|---|---|
| 0–19 points — engines broadly agree | 40 |
| 20–39 points — partial disagreement | 4 |
| 40+ points — engines fundamentally disagree | 46 |
Only 4 of 90 subjects sit in the middle band. A brand is either seen roughly the same way everywhere, or it is seen completely differently depending on which engine you ask — there is no gradual middle. Median gap: 46.8 points. Largest gap measured: 91.5 points, one engine near-complete and another near-zero on the same subject on the same day.
- 63% of subjects were recognised by all four engines
- 9% were recognised by none of them
- The remaining 28% exist in some engines and not others
For that 28%, the question “am I visible in AI?” has no single answer. It depends entirely on which AI you ask.
Finding 4 — websites are not blocking AI crawlers; something else is missing
In the technical audit of the 100 websites in this sample:
- 0 of 100 blocked AI search crawlers. 0 of 100 blocked AI training crawlers. Not a low number — zero.
- 45 of 100 had no robots.txt at all
- 70 of 100 had no llms.txt
- 71 of 100 had no discoverable sitemap
Contrary to the common assumption, sites are not shutting AI crawlers out — in this sample not one of them did. Access is not the constraint. What is missing is machine-readable structure worth citing: a page an engine can parse into a claim it is willing to repeat. The gap is not a locked door, it is an empty room.
Finding 5 — a single measurement is not a measurement
This is the finding we would most like other people to test against their own data, because it changes how every number above should be read.
Engine answers are non-deterministic. In repeat runs of the same subject on the same day, we observed recognition flipping between runs: an engine recognises the subject, then does not, then does again, within a window of well under two hours. Because recognition gates a large share of any visibility score, a single flip moves the headline number dramatically.
The practical consequence for anyone building or buying this kind of measurement: one reading is noise. What makes a number trustworthy is repeat sampling, a dated log of every run and the prompt that produced it, and smoothing across runs rather than reporting the latest value. A tool that shows you one number with no history is showing you a coin flip.
Method and limits
What we did not measure matters as much as what we did:
- Sample: 154 unique subjects (100 websites, 29 social accounts, 17 people, 8 brands). Where the same subject was scanned more than once, only the most recent measurement counts. 90 of them carry a complete four-engine measurement; the engine-comparison findings use those 90.
- Window: 9 July – 24 August 2026. Model versions change; these numbers are a photograph, not a fixed truth.
- Language and geography bias: the sample skews towards Turkish-language and Turkey-based subjects. This is not a global average and should not be read as one.
- Judge: the accuracy score is assigned by a model different from the ones that produced the answers, and the median of three measurements is used, to reduce self-preference bias.
- “Hallucination suspicion” is not a claim of proven error. It flags assertions the judge model could not verify against live web data, including plausible ones. Anyone comparing this figure with their own should first align on that definition — the number is meaningless across two different definitions.
- Grok was excluded. It appeared in earlier measurements in shadow mode and was later switched off entirely. Including it would create a false trend in any follow-up comparison.
What to do with this
The practical takeaway in one sentence: visibility cannot be measured on one engine, recognition is not accuracy, and one reading is not a measurement. If you are checking your own brand, check it on all four, check what sources the answers cite, and check it more than once before believing the number.
Aggregated from GEONI's production scan database on 25 August 2026. Quoting is welcome; a citation of geoni.ai/guides/ai-visibility-measurement-report is sufficient. If you run a comparable measurement and get different results, we would genuinely like to see it — the definitions matter more than the numbers.
Where does your own brand stand across four engines?
GEONI asks your category questions to ChatGPT, Claude, Gemini and Perplexity separately, extracts the sources those answers cite, checks the Google AI Overview box, and verifies that AI crawlers can reach your site. Your first scan is free.
Run a free scan →