Why does every AI visibility tool give you a different score?
AI visibility tools disagree because AI engines themselves are not deterministic, and every tool samples them differently. Different prompts, different engines, different run counts, different scoring weights, different days. Each tool is photographing a moving target from a.
AI visibility tools disagree because AI engines themselves are not deterministic, and every tool samples them differently. Different prompts, different engines, different run counts, different scoring weights, different days. Each tool is photographing a moving target from a different angle, then presenting its photo as the territory. See generative engine optimization for the wider picture.
This piece explains where the differences come from, when a score is worth trusting, and how to read any AI visibility number like someone who knows how it was made.
What an AI visibility tool actually measures
An AI visibility tool asks AI engines questions a buyer might ask, records whether your brand appears in the answers, and compresses those observations into a score. That is the whole product. Everything that matters lives in the details: which questions, which engines, how many times, and how the observations become a number.
There is no registry of AI answers to look up. Every datapoint has to be generated by actually asking the engine, at a moment in time, with a specific phrasing. The tools are not reading a scoreboard. They are running experiments.
Five reasons the scores never match
Every divergence between two tools traces back to at least one of these five choices. None of them are visible on the score itself.
AI answers are not deterministic
Ask ChatGPT the same question twice and you can get different answers with different sources cited. This is not a bug in the tools; it is how large language models work: outputs are sampled, not retrieved, as OpenAI's own text generation documentation describes. A tool that queries once and reports a precise-looking number has measured a coin flip to two decimal places.
Every tool asks different questions
One tool audits "best crypto exchange." Another audits "safest exchange for beginners in Europe." Real buyers ask the second kind. Whether a tool's prompt set matches how your buyers actually phrase things changes the result more than anything your marketing team did last quarter.
Every tool samples different engines
ChatGPT, Perplexity, Gemini, Claude and Google AI Mode disagree with each other constantly, and Google's AI surfaces follow their own rules entirely, documented in Google's AI features guidance for Search. In our own audits, the same brand routinely gets a strong showing on one engine and a blank on another. A tool that averages three engines will never match a tool that samples five, and neither is measuring "AI" as a whole.
Every tool weights differently
Is being mentioned worth half a citation? Does position in the answer matter? Is Perplexity worth as much as ChatGPT? These are editorial decisions, not measurements. Two tools can record identical observations and still publish different scores, because the weighting is opinion. The honest ones publish those opinions. Most do not.
The answers drift over time
Engines update, indexes refresh, sources rise and fall. The AEO research community keeps documenting how fast this ground moves; Ahrefs' published studies on AI search behavior are a good running record. A score from three weeks ago describes a system that no longer exists. Two tools that scanned on different days measured genuinely different realities.
The five choices, and what each one moves
None of these five appear on the score itself. Each one alone can move a brand's number by more than a quarter's worth of marketing work.
| Choice | What varies between tools | What citeOS does |
|---|---|---|
| Non-determinism | One probe per query, or many | Multi-sample per query, majority-merged, per-engine stability reported |
| Prompt set | Head keywords vs real buyer phrasing | 20 buyer-style prompts per category, editable prompt library |
| Engine coverage | Three engines, or five, or one | 5 engines, 100 observation points |
| Scoring weights | Usually undisclosed | Mention 0.60, citation 0.20, authority 0.20. All published |
| Drift over time | Rescored silently, or not at all | Weekly refresh, dated, with prior runs kept |
A tool is not lying because its number differs from another tool's. It is lying if it will not tell you which of these five it chose. That is the only question worth asking a vendor.
So are the tools lying?
Mostly no. The dishonesty is rarely in the measurement; it is in the precision theater. Reporting "your AI visibility is 34.7" implies a stability the underlying system does not have.
The tell is what a tool discloses. If you cannot find the prompt list, the engine list, the sample count, and the scoring weights, you are not looking at a measurement. You are looking at marketing that outputs digits.
How to read any AI visibility score
Which prompts?
If you cannot see the questions, the score is unanchored.
Which engines, listed by name?
"AI search" is not an engine. A real answer names ChatGPT, Perplexity, Gemini, Claude, Google AI Mode.
How many samples per prompt?
One run is an anecdote. Multiple runs with a stability measure is data.
What are the weights?
If they are secret, the score is an opinion wearing a lab coat.
When was it measured?
A score without a date describes nothing.
A tool that answers all five can be compared against itself over time, which is the only comparison that matters. Chasing score parity between two different tools is chasing an artifact of their methodologies.
Want a score you can actually interrogate? Run the disclosure test on us.
Scan my brand free →What honest measurement looks like
We build citeOS around the disclosure test because our buyers are crypto teams, and crypto buyers are professionally allergic to numbers nobody can verify.
A citeOS audit runs 20 buyer-style prompts across all five major engines, which produces 100 observation points per scan. Paid audits sample each prompt multiple times and report stability, not false precision. Every scoring weight is published on our methodology page, and every citation we count can be reproduced by asking the engine the same question yourself. When an engine fails or returns no AI surface, we show that honestly instead of a fabricated zero.
Not because that makes prettier numbers. Because the score is the beginning of the work, not the product. The product is knowing which sources moved the answer, and what to ship next.
AI visibility scores, answered.
What is an AI visibility tool?
Software that measures whether AI engines like ChatGPT, Perplexity and Gemini mention or recommend your brand when users ask buying questions in your category. It works by querying the engines directly and recording the answers.
Which AI visibility score is the most accurate?
None of them are "accurate" in an absolute sense, because AI answers change between runs. The useful question is which tool is most transparent: published prompts, named engines, multi-sample runs, and public scoring weights let you trust a score's direction over time.
Why did my AI visibility score change when nothing changed on my site?
Because the engines changed. Model updates, index refreshes and source shifts move answers constantly. This is why single-run scores are unreliable and trend lines beat snapshots.
Can I improve my AI visibility score?
Yes, but not by optimizing for the score. AI engines cite sources they trust: independent media coverage, community presence, and content that answers buyer questions directly. Improve those and every honest tool's score follows. Start by seeing where you stand with a free AI citation scan.
Go one level deeper.
Get a score you can interrogate.
Scan my brand free →