If you run your brand through two AI visibility tools in the same week and get 34% from one and 61% from the other, neither is lying. They show different numbers because each picks its own prompt set, its own denominator (out of all answers, or only branded ones), its own run count, its own engines and regions, and its own definition of “appearing.” Any one of those choices can move the score by tens of points with nothing about your brand changing. So the practical rule is: an AI visibility score is only meaningful inside one tool over time, never compared across tools. This guide explains exactly why the numbers diverge, and how to run an honest comparison anyway, by comparing behavior instead of scores.
The trap is assuming an AI visibility score behaves like a Google ranking, a stable number two people can check and roughly agree on. It does not. It is a photograph of a moving target, taken through a lens each vendor grinds differently.
”50%” can mean four different things

Start with the single biggest source of divergence: the denominator. A score of 50 can mean you appeared in half of all answers the tool checked, brand-free ones included. Or half of the answers that named any brand at all, a much smaller pool. Or an impression-weighted share of a competitive set, weighted by how prominent each mention was. Or a blended composite mixing coverage, consistency, and citations into one made-up index. Those are four different measurements wearing the same “50%” label.
This matters more than it sounds, because most AI answers name no brand at all. Analyses in 2026 found a majority of answers are brand-free, so switching from “out of all answers” to “out of branded answers” can roughly double the score by itself. Nothing about your brand changed. Only the definition of “out of what” did. That is why two tools reporting 34% and 61% are usually not disagreeing, they are answering two different questions. The full breakdown of these metric types is in mention vs rank vs visibility vs citation.
Five reasons two tools disagree

Five knobs, all set differently, all invisible on the dashboard. First, different prompt sets: each tool picks its own questions or lets you supply yours, and different questions produce a different score before anything else happens. Second, different denominators: out of all answers, or only branded ones, the single choice that can nearly double or halve the same underlying result. Third, different run counts and timing: answers vary run to run and drift over time, so a tool sampling once and a tool sampling thirty times are looking at different worlds, the math of which is in how many runs a number needs before you can trust it. Fourth, different engines, regions, and login state: cited sources barely overlap across engines, with research finding only a tiny fraction of sources appear on all major engines and most appearing on just one, so a US logged-out check and a UK logged-in one are simply not the same measurement. Fifth, different definitions of “appearing”: a plain mention, a link, or a link in first position, each tool draws the line somewhere and rarely in the same place.
Set all five knobs differently, which every vendor does, and identical underlying reality produces wildly different headline numbers. The dashboard shows you the number and hides all five settings behind it.
The moving target underneath it all
Even if two tools matched all five settings, they would still diverge, because the thing being measured moves. LLM answers are non-deterministic: one clean public demonstration sampled the same prompt a thousand times at temperature zero, the setting meant to force identical output, and got dozens of different completions. On top of that, the top-ranked brand for a given prompt flips more than half the time between runs on some engines, and citation volumes for the same brand can differ by hundreds of times across platforms. So any single score is, as one analysis put it, a photograph with a date on it. A number from May quoted in an August report without the date is not a measurement, it is a fossil. This underlying variance is covered in which ChatGPT you are actually tracking.
The honest consequence: the industry itself admits the gap. Surveys in 2026 found that a large majority of marketing leaders cannot accurately measure their AI visibility, and only a tiny fraction believe they have the tools to track every relevant metric across platforms. The demand exploded (searches for these tools rose more than tenfold in a year) faster than any shared standard could form.
How to run an honest comparison

You cannot compare their numbers, but you can compare their behavior. First, give both tools the identical prompt list, the same 20 to 30 real buyer questions in each, and note that if a tool will not let you supply prompts, that itself tells you something. Second, ignore the two absolute scores entirely: one says 34, one says 61, and trying to reconcile them is wasted effort because they answer different questions. Third, compare the ranked list of competitors instead: do both tools agree on who owns the category and where you sit relative to rivals, because order is comparable even when the score is not. Fourth, pick one and only ever compare it to itself, because a consistent method run the same way each week is what makes a trend real, and you should never compare across tools or across months.
The single rule that survives all of this: the absolute number is only meaningful inside one tool over time. Between tools, watch the ranking, not the score. And demand the two things every honest vendor should show you, the prompt list and the denominator, because a tool that hides both is giving you a rating, not a measurement.
What to look for in a tool
Given all this, the tools worth trusting are the ones that are explicit rather than impressive. They let you set the prompt list, they tell you the denominator, they sample each prompt enough times to beat the noise, they split results by engine instead of blending, and they show the raw answers behind the score so you can audit it. Rankry is built on that principle: your prompts, a stated methodology, per-engine results across ChatGPT, Claude, Gemini, Perplexity, and Grok, and the actual answers and cited sources kept as evidence, from $99 a month on a no-card trial. It will not promise that its 47% matches someone else’s 47%, because no honest tool can. It promises that its 47% means the same thing next week, which is the only comparison that was ever real. The broader loop is in how to monitor your brand across AI search engines.
FAQ
Why do two AI visibility tools give me different scores? Because they measure different things. Each picks its own prompt set, denominator, run count, engines, regions, and definition of “appearing,” and most do not publish any of it. Two scores from two tools are two answers to two different questions.
Can an AI visibility score be compared between tools? No. A 50 from one tool and a 50 from another can mean completely different things, half of all answers versus half of branded answers versus a blended index. Compare a score only against itself over time, within one tool.
Why does changing the denominator change my score so much? Because most AI answers name no brand at all. Dividing by only the answers that mention some brand, instead of by all answers, shrinks the pool and can roughly double the score, with nothing about your brand having changed.
Are AI visibility scores accurate at all? They are useful as a consistent trend inside one tool, not as an absolute truth. The number is often reported with more precision than it has. Widen the sample, split by engine, and read the number as a range, not a fact.
How should I compare two AI visibility tools? Give both the identical prompt list, ignore the absolute scores, and compare the ranked competitor lists instead, order is comparable even when the score is not. Then pick one tool and only ever compare it to itself over time.
What makes an AI visibility tool trustworthy? Transparency over polish: it lets you supply prompts, states its denominator, samples each prompt many times, separates results by engine, and shows the raw answers behind the score. A tool that hides its prompt list and formula is giving you a rating, not a measurement.
Your prompts, a stated method, per-engine results, and the raw answers behind every score, the same meaning every week. Start a free 7-day Rankry trial, no card, first report in two minutes.