WWikiAP
Category: Methodology

Why AI Visibility Tools Report Different Scores for the Same Brand: Causes and a Verification Procedure

Why the same brand scores differently in every AI visibility tool, split into four layers (engine variance, collection design, metric definition, aggregation), with a side-by-side table of the published formulas from Profound, Peec AI, and Ahrefs, sample-size confidence interval thresholds, and a six-step verification procedure.

Technical GEO 에디터Published

Measure the same brand in the same week with two tools and the scores will not match. The first reaction is usually "which tool is right," and that question goes nowhere. The two numbers came from different prompt sets, run a different number of times, divided by different denominators. This page splits the reason AI visibility tools disagree into four layers, engine, collection, definition, and aggregation, then walks through the verification steps a dashboard number should clear before it lands in a report.

Four layers behind the gap, defined in 30 seconds

Engine variance is the tendency of an AI engine to return a different answer, and different sources, on each run of the same prompt.

Collection design is the rule set that decides which prompts get sent, how many times, and from which region and login state.

Metric definition is the counting rule that fixes what counts as one appearance and what goes into the denominator.

Aggregation is the weighting and averaging method that folds per-prompt and per-engine results into a single score.

The four choices multiply. Two tools that decide each layer a little differently will report different numbers for the same brand. Visibility as a design variable rather than a fixed constant was already an assumption in the research that first formalized GEO. That paper defined visibility metrics inside generative engine responses, proposed a black-box optimization framework, and reported visibility gains of up to 40%.[13]

Four layers behind the score gap 1. Engine variance Model version, nondeterminism Region, login state 2. Collection design Prompt list, sources Run count, cadence 3. Metric definition Mention or citation Denominator, position 4. Aggregation Engine weights, averages Competitor set Tool A reported score e.g. 18% mention rate Tool B reported score e.g. 30% mention rate Same brand, same week, different numbers
Four layers behind the score gap, where differing choices from engine variance through aggregation end up as a gap in the final score. The two values at the bottom are a hypothetical used to illustrate the denominator effect.

No engine answers the same question twice

The first layer sits outside any tool's control. SE Ranking pushed the same 10,000 keywords through Google AI Mode three times on one day and compared 9,734 responses carrying 122,617 links. Exact URL overlap across the three datasets averaged 9.2%, and 2,006 keywords (21.2%) shared no URL at all across the three runs.[1]

Google AI Mode, result overlap across 3 same-day runs URL, all 3 runs 9.2% Domain, all 3 runs 14.7% URL, 2-run pairs 18.5~19% Domain, 2-run pairs 26.6~27% Source: (SE Ranking, 2025)
Google AI Mode, result overlap across three same-day runs of the same keyword, source: (SE Ranking, 2025). The 2-run pair figures differ across the three dataset pairs, so they are shown as a range and the bar length is the midpoint of that range, the per-pair values are in the table below.
Match criterionAll 3 runsDatasets 1 and 2Datasets 1 and 3Datasets 2 and 3
Exact URL match9.2%18.5%18.5%19%
Domain match14.7%26.7%26.6%27%

Every figure comes from (SE Ranking, 2025). The per-pair values scatter between 18.5% and 19% for URLs and between 26.6% and 27% for domains, which is the range the bar chart above reports. Any two answer sets shared 3.3 identical URLs and 3.7 identical domains on average.[1]

A single response carried 12.6 links on average.[1] That is roughly a dozen citation slots, redrawn almost from scratch on every run. At the brand level the swing shows up even more plainly. SparkToro compiled 2,961 runs of 12 prompts across ChatGPT, Claude, and Google AI answers by 600 volunteers, and reported that two runs return the same brand list less than once in 100 attempts, with the same list in the same order closer to once in 1,000 (collected November and December 2025).[3] A separate experiment ran 12 prompts 100 times each on logged-out free ChatGPT, 1,200 runs in total from different IP addresses, and found roughly 44 brands mentioned across the 100 responses per prompt, of which only about 5, near 11% of the total, appeared in more than 80% of those responses (Search Engine Land, 2026).[4]

The cause sits in inference infrastructure, not service policy. Thinking Machines Lab sampled one prompt 1,000 times at temperature 0 and got 80 distinct completions; the first 102 tokens matched, then the 103rd token split. The main driver was not floating-point nondeterminism by itself but kernel operations whose results change with batch size, and swapping in batch-invariant kernels made all 1,000 runs identical.[2] API documentation says the same thing. OpenAI states that setting a seed does not guarantee determinism and that backend configuration changes should be monitored through system_fingerprint.[10]

Platform differences are documented too. Google explains that AI Mode and AI Overviews can run on different models and techniques, so the answers and link sets they surface diverge.[9] One number labeled "Google AI visibility" cannot be interpreted without knowing which surface it measured. Multi-engine measurement methodology covers the per-engine approach in more detail.

Every tool divides by a different denominator

The second cause is the formula. Put the public documentation side by side and the same metric name turns out to describe different arithmetic.

ToolHeadline metricDenominator and calculation basis (per public docs)Prompts and collection designMethodology disclosure
ProfoundVisibility, Share of VoiceVisibility divides responses containing the brand by responses containing at least one brand; SoV divides own-brand mentions by total brand mentions[5]Registered prompts sent automatically to answer engines on a daily cycle[5]Numerator and denominator of both metrics stated in the docs
Peec AIVisibility ScoreResponses mentioning the brand divided by total responses, multiplied by 100[6]Applied across all tracked prompts and competitor targets[6]Formula published; mention-detection criteria and weighting absent from the docs
Ahrefs Brand RadarMentions, Citations, Impressions, AI Share of VoiceMentions counted once per response, citations once per domain; impressions sum the search volume of prompts where the brand appears; AI SoV is impression share, weighted by impressions across multiple platforms[7]Search-derived prompts; impressions use the highest-volume keyword that surfaces the prompt in People Also Ask[7]Counting units and weighting method documented, with a note that values shift substantially with entity settings
Semrush AI VisibilityShare of Voice, Visibility, MentionsSoV is mention share against competitors; Visibility scores top citation position across tracked prompts, so holding the first citation on every prompt equals 100%; Mentions counts prompts containing the brand[8]Prompts registered in a tracking campaign, collected against ChatGPT and Google AI Mode[8]Definitions published; the mathematical formula and prompt selection method sit in separate documentation
BOIDA BVIBVI (Brand Visibility Index)Claims to measure brand exposure in major generative AI answers across several dimensions and convert it into an index; formula and weights are not published[14]Described as analyzing exposure across major generative AI from a URL or product name alone (AVING, 2026)[15]Index formula not published; the named engine list and collection cadence have to be confirmed against official material

Price is left out of the comparison axis. What splits the numbers between tools is the formula and the collection design, not the plan tier, and those are also what decide whether a result can be reproduced in practice. Public pricing moves, so check each vendor's official page during evaluation. Price bands are collected in AI visibility monitoring tool comparison.

The size of the denominator effect is easy to check with arithmetic. Say 60 of 100 responses name at least one brand, and 18 of those name yours. Divide by all responses and the score reads 18%. Divide by responses that name at least one brand and it reads 30%. Same data, one definitional change, 12pp apart. This example is a hypothetical constructed to show the denominator effect, not output from any specific tool.

One more axis matters. A brand mention and a domain citation are different events. Your name can appear in an answer while your domain goes uncited, and your page can be used as a source while the brand name never shows up in the sentence. That is why Ahrefs documentation separates mentions and citations into distinct metrics and states that citations count only pages shown as inline sources in the response, excluding pages that were crawled but not cited.[7] Align the metric names without aligning the event definitions and the two tools will keep producing different answers. Ahrefs Brand Radar vs Semrush AI Toolkit sets the two products' metric designs side by side. AI search share of voice picks up how share metrics are structured.

With a small sample, most of the gap is noise

The third cause is the sample. A visibility score is a proportion estimate, so a small sample means a wide confidence interval. Compute the 95% interval for a binomial proportion with the normal approximation and the width follows from the observed rate and the sample size. The figures in the table below are not any tool's reported output, they are computed directly from that formula. The NIST/SEMATECH handbook notes that this approximation is the most commonly used one but calls it inferior, because its lower limit can drop to an impossible value, and recommends the Wilson method instead, so treat the table as an order-of-magnitude guide.[12]

Response sample (n)95% CI half-width at an observed 30% mention rateSmallest gap you can call
30about ±16.4ppLarge gaps only
50about ±12.7ppLarge gaps only
100about ±9.0ppDouble-digit pp gaps
300about ±5.2ppHigh single-digit pp gaps
1,000about ±2.8ppGaps of a few pp

Response sample size here is prompts times runs times engines. Fifty prompts run once on one engine gives a sample of 50, and at that size the move from 28% last week to 33% this week supports no claim about performance. GEO KPI measurement goes further into repeat measurement and prompt panel design.

Grading measurement quality has already drawn a standardization attempt. In August 2026, IAB's "Measuring Visibility in the AI Era" split AI visibility data into directional and decision-grade, and set out the principle that measurement vendors disclose their query selection method, collection cadence, and aggregation methodology.[11] The full framework is explained in IAB Four Ps measurement standard.

Verification, in six steps

None of this argues for throwing the numbers out. It argues for drawing a line around how far they can be trusted. The sequence starts by pinning down the spec.

StepWhat to doPass test
1. Pin the specPut the prompt list, engines and models, region, language, login state, run count, and collection cadence in writingCould a third party reproduce these conditions?
2. Align definitionsSeparate brand mentions from domain citations, then record each metric's numerator and denominator from the tool's own documentationDo the two tools share a denominator, and if not, is the difference written down?
3. Size the sampleSet the gap you need to call, back out the required response count with the normal approximation, then fix prompt count and run countIs the confidence interval narrower than the gap you want to call?
4. ReproduceRe-run the same prompt set manually or through a separate script, log whether the brand appears, and compare against the tool's reported figureDo the confidence intervals of the two estimates overlap?
5. Isolate the layerIf a gap remains, match conditions one at a time in the order engine, collection, definition, aggregationCan you name the layer the gap comes from?
6. Track changesAnnotate model version swaps, tool methodology updates, and prompt set changes, and break the trend line at those pointsIs there an annotation at each sharp turn in the trend?

Steps 1 and 2 are the ones teams skip most often. Without a spec, nobody can later reconstruct the conditions a number came from. Without definitions, tool-to-tool gaps keep getting mistaken for arguments about performance.

Step 4 needs controlled conditions. Make logged out, fixed region, fresh session, and no personalization history the defaults, and record the conditions used. SparkToro disclosed that its study did not align temperature, country, device, or chat history and used whatever default experience each participant already had, and that disclosure is exactly the kind of information that changes how a result reads.[3]

Step 5 runs faster in order. Match engine and model first, then prompt set and run count, then the denominator. Whatever survives those three often traces back to aggregation weights. Merging multiple platforms with an impression-weighted average, for example, lets high-volume prompts dominate the score.[7]

Verifying visibility from log data confuses the layers. Google states that performance for sites surfaced in AI features is folded into the Web type within total Search Console search traffic and is not broken out separately.[9] Traffic and conversion data describe what happens after a citation, so manage them as a different metric from in-answer mention rate.

What else splits in Korean-language measurement

Korean brands carry one more variable: query language and the tracked engine set. The same brand draws a different answer mix from English and Korean queries, and when tools track different engine lists, the denominator itself changes. Global tools work well for tracking the major overseas engines, while coverage of domestic Korean engines varies by product, so confirm it against official material before buying.

Domestic options are growing. BOIDA, run by Designovel, states that its BVI metric measures brand exposure in major generative AI answers and carries that through to remediation as part of its product scope[14]. Press coverage describes it as analyzing exposure across major generative AI from a URL or product name alone (AVING, 2026)[15]. The named engine list and the way Korean-language queries are handled are best confirmed against official material. Other domestic vendors include Next-T, Across, LeadGenLab, and Ascent AI; what each one publishes, and where its free tier ends, is laid out spec by spec in the comparison of free GEO audit tools in Korea. Ranking or first-place phrasing that Korean vendors apply to themselves is subject to fact-checking, so read it as claimed positioning and compare instead on documented engine lists, query selection methods, and disclosure of collection cadence.

Takeaways

The gap between tools is not a bug. Engine answers change from run to run, tools send different prompts different numbers of times, and metrics sharing a name divide by different denominators, so the gap is a structural outcome. The 9.2% exact URL overlap across three same-day Google AI Mode runs, on its own, sets the resolution of any single measurement.[1]

Three conclusions carry into practice. First, do not compare absolute values across tools. Read only the trend and the relative competitor position inside one tool under one spec. Second, attach the spec and the confidence interval to any number that goes into a report; one line listing prompt count, run count, engines, period, and denominator definition makes next quarter's reading far easier. Third, put methodology disclosure into your tool selection criteria. A product that publishes its formula and collection design is a product you can verify. What is GEO covers the definition and scope, and AI visibility monitoring tool comparison continues on tool choice.

Related companies

Frequently asked questions

Q.AI visibility tools disagree on my score. Which one is correct?
When the denominators and the collection designs differ, both numbers are correct inside their own definitions. Profound's visibility score divides by responses that contain at least one brand, while Peec AI divides by all responses. Ahrefs uses an impression share converted from search volume rather than a count of mentions. Absolute values therefore do not compare across tools. What compares is the trend inside one tool on one prompt set, plus your position relative to competitors.
Q.Why does my score move week to week even on the same tool?
Because engine answers are nondeterministic. Running the same keyword three times on the same day in Google AI Mode produced an average exact URL overlap of 9.2%, and 21.2% of keywords shared no URL at all across the three runs (SE Ranking, 2025). Model version swaps, region, login state, and sampling error stack on top of that. When the movement sits inside the confidence interval, read it as noise rather than a change in performance.
Q.How many prompts, and how many runs each?
Start from the precision you need. At an observed 30% mention rate, the 95% confidence interval runs about ±12.7pp on 50 responses, about ±9.0pp on 100, about ±5.2pp on 300, and about ±2.8pp on 1,000 (computed with the normal approximation in the NIST/SEMATECH handbook, which recommends the Wilson method over this approximation because its lower limit can fall below zero). Response count here is prompts times runs times engines. To call a gap of a given size against a competitor, the interval has to be narrower than that gap.
Q.Can I verify a tool's score by hand?
Partly. Run the same prompts several times while logged out, with region fixed and history cleared, log whether the brand appears, and check whether your estimate's confidence interval overlaps the mention rate the tool reported. A test that leaves personalization and memory settings uncontrolled will carry all the variation of a real user environment. The SparkToro study likewise did not align country, device, or chat history across participants, and ran on whatever default experience each person already had (SparkToro, 2026).
Q.Can Search Console or GA4 verify a tool's score?
Not visibility itself. Google states that data for sites surfaced in AI features is included in the Web type within total Search Console search traffic and is not broken out separately (Google Search Central). Logs and referrer data describe traffic that arrives after a citation, so they cannot stand in for in-answer mention rate. Treat the two as metrics on different layers and manage them separately.
Q.What extra checks apply to Korean-language measurement?
Query language, the tracked engine list, and whether domestic Korean engines are covered. The same brand draws a different answer mix from English and Korean queries, and when the tracked engine set differs, so does the denominator. During tool selection, ask for documentation on Korean query support, a named list of tracked engines, and disclosure of query selection and collection cadence, then compare.

Sources

  1. [1] ↑AI Mode Research: Sources, Volatility, and Differences between AIO and Organic SearchSE Ranking
  2. [2] ↑Defeating Nondeterminism in LLM InferenceThinking Machines Lab
  3. [3] ↑AIs are highly inconsistent when recommending brands or productsSparkToro
  4. [4] ↑What repeated ChatGPT runs reveal about brand visibilitySearch Engine Land
  5. [5] ↑Answer Engine Insights OverviewProfound
  6. [6] ↑Visibility (Brand metrics)Peec AI Docs
  7. [7] ↑AI Visibility MetricsAhrefs Help Center
  8. [8] ↑AI Visibility MetricsSemrush Knowledge Base
  9. [9] ↑AI features and your websiteGoogle Search Central
  10. [10] ↑Reproducible outputs with the seed parameterOpenAI Cookbook
  11. [11] ↑Measuring Visibility in the AI EraIAB
  12. [12] ↑7.2.4.1. Confidence intervals (proportion)NIST/SEMATECH e-Handbook of Statistical Methods
  13. [13] ↑GEO: Generative Engine OptimizationarXiv / KDD 2024
  14. [14] ↑BOIDA (BVI) 서비스Designovel
  15. [15] ↑디자이노블, 생성형 AI 브랜드 가시성 지표 BVI 공개AVING News

This document was last edited on Sep 9, 2026. WikiAP content is compiled from public primary sources and updated for accuracy.