// GEONI Measurement Report

How Four AI Engines See the Same Brand

25 August 2026 · GEONI Team · Measurement window: 9 July – 24 August 2026

Aggregated results from 154 unique measured subjects — websites, brands, people and social accounts — scanned across ChatGPT, Claude, Gemini and Perplexity. Every number on this page comes from our own measurement. There are no industry estimates and no figures quoted from anyone else. No customer name, domain or individual score appears anywhere.

Finding 1 — the most used engine recognises the fewest brands

Engine Recognised the subject Hallucination suspicion Judge accuracy score Average visibility
Perplexity89%19%75.265.6
Claude84%13%76.059.7
Gemini90%29%68.659.2
ChatGPT67%12%72.945.7

The engine with the largest user base recognised the fewest subjects. ChatGPT knew 67% of what we measured; Gemini knew 90%. The practical consequence: being known to Gemini tells you nothing about being known to ChatGPT. A brand that checks its visibility in one chat window and concludes it is fine has measured one quarter of the surface.

Finding 2 — the engine that recognises most also invents most

Every answer was compared against live web data by an independent judge model looking for contradictions and unverifiable claims. Gemini carried hallucination suspicion in 29% of measurements — roughly 2.4× ChatGPT's 12%. Gemini also had the lowest judge accuracy score at 68.6, while Claude had the highest at 76.0.

The conclusion runs opposite to intuition: being recognised and being described correctly are different properties. An engine mentioning your brand is not evidence that it is saying true things about you, and the engine most likely to mention you is the one most likely to be wrong.

Finding 3 — engines either agree closely or disagree wildly, with almost nothing in between

For each subject measured on all four engines we took the gap between the highest and lowest engine score. We expected a normal-looking spread around some average. What we found was bimodal:

Gap between best and worst engine Subjects
0–19 points — engines broadly agree40
20–39 points — partial disagreement4
40+ points — engines fundamentally disagree46

Only 4 of 90 subjects sit in the middle band. A brand is either seen roughly the same way everywhere, or it is seen completely differently depending on which engine you ask — there is no gradual middle. Median gap: 46.8 points. Largest gap measured: 91.5 points, one engine near-complete and another near-zero on the same subject on the same day.

For that 28%, the question “am I visible in AI?” has no single answer. It depends entirely on which AI you ask.

Finding 4 — websites are not blocking AI crawlers; something else is missing

In the technical audit of the 100 websites in this sample:

Contrary to the common assumption, sites are not shutting AI crawlers out — in this sample not one of them did. Access is not the constraint. What is missing is machine-readable structure worth citing: a page an engine can parse into a claim it is willing to repeat. The gap is not a locked door, it is an empty room.

Finding 5 — a single measurement is not a measurement

This is the finding we would most like other people to test against their own data, because it changes how every number above should be read.

Engine answers are non-deterministic. In repeat runs of the same subject on the same day, we observed recognition flipping between runs: an engine recognises the subject, then does not, then does again, within a window of well under two hours. Because recognition gates a large share of any visibility score, a single flip moves the headline number dramatically.

The practical consequence for anyone building or buying this kind of measurement: one reading is noise. What makes a number trustworthy is repeat sampling, a dated log of every run and the prompt that produced it, and smoothing across runs rather than reporting the latest value. A tool that shows you one number with no history is showing you a coin flip.

Method and limits

What we did not measure matters as much as what we did:

What to do with this

The practical takeaway in one sentence: visibility cannot be measured on one engine, recognition is not accuracy, and one reading is not a measurement. If you are checking your own brand, check it on all four, check what sources the answers cite, and check it more than once before believing the number.

Aggregated from GEONI's production scan database on 25 August 2026. Quoting is welcome; a citation of geoni.ai/guides/ai-visibility-measurement-report is sufficient. If you run a comparable measurement and get different results, we would genuinely like to see it — the definitions matter more than the numbers.

Where does your own brand stand across four engines?

GEONI asks your category questions to ChatGPT, Claude, Gemini and Perplexity separately, extracts the sources those answers cite, checks the Google AI Overview box, and verifies that AI crawlers can reach your site. Your first scan is free.

Run a free scan →

Other guides