Why AI Visibility Scores Differ Between Tools
AI visibility scores differ because tools choose different engines, prompts, denominators, source roles, evidence floors, and freshness windows. A practical Para Labs guide for comparing scorecards without treating them as interchangeable truth.
AI visibility scores differ between tools because each platform chooses its own engines, prompt set, denominator, source-role rules, evidence floor, and freshness window. A score is useful only when the buyer can see what was measured, what was excluded, and whether the result is a stable citation rate or a thin observation. (Machine Relations Index public data)
Key takeaways
- AI visibility tools can disagree without any tool being broken; they may be measuring different answer surfaces, dates, prompts, geographies, and source roles.
- The September 15 Machine Relations Index v2 release reports 122,144 citation events, 15,468 observed answer runs, 21,861 domains, and 82 published strata across six healthy engines.
- MRI v2 publishes citation rates only after a segment clears at least 10 observations across at least seven distinct run dates, which is a methodological choice other tools may not share.
- CMOs should compare scorecards by method, not by headline number.
- The practical operating move is to separate detection, diagnosis, and correction before buying or renewing an AI visibility platform.
AI visibility scores differ because the measurement object differs
An AI visibility score is not a universal market share number; it is a model of a measured question set. One tool may ask brand prompts, another may ask category prompts, and a third may measure whether sources were cited in buying questions. Those are adjacent problems, not the same denominator.
The Machine Relations Index v2 makes this explicit. Its public methodology measures how often AI answer engines cite source domains across category-question strata, not how often a named brand is mentioned in every possible prompt. The September 15 release covers six answer engines: ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, and Perplexity. (MRI release manifest)
That method is useful because it ties visibility to answer-engine source selection. It also shows why another dashboard can produce a different result if it measures brand mentions, sentiment, rank position, answer inclusion, referral traffic, or screenshots instead of domain citation rates.
Engine coverage changes every AI visibility score
A scorecard that measures six engines is not comparable to a scorecard that measures one model or one search surface. The September 15 MRI release says all six engines were healthy within the observation window. That matters because engines do not cite the same sources at the same rate, and some answer surfaces retrieve web documents differently from pure chatbot sessions. (MRI release manifest)
Google's own guidance for generative AI features keeps the work tied to retrievable, reliable public content rather than a separate magic optimization layer. Google says site owners should keep following Search fundamentals for helpful, reliable, people-first content while AI features may use indexed web sources. (Google Search Central)
For buyers, this means an AI visibility score should always disclose the engine roster. A ChatGPT-only score can be directionally useful, but it cannot be treated as the answer for Google AI Mode, Perplexity, Claude, Gemini, or AI Overviews.
Prompt baskets change AI visibility denominators
Prompt design determines what the score can mean. A brand-defense prompt such as "is Brand X trustworthy" answers a different business question from a buyer prompt such as "best vendor for enterprise data governance" or a source-selection prompt such as "which publications do AI engines cite for B2B software buying questions."
The Machine Relations Index groups observations into strata: a subject category paired with a question type. The September 15 public release reports 157 total strata, 151 collectable strata, 82 published strata, 69 collecting strata, and 51 empty strata. Those fields prevent readers from treating every topic as equally mature. (Machine Relations Index public data)
A visibility vendor that reports one blended score across brand, category, competitor, and news prompts may be useful for an executive dashboard. It is weaker for diagnosis unless the underlying prompt basket is visible.
Evidence floors change whether thin signal becomes a score
A stricter evidence floor produces fewer but more interpretable scores. MRI v2 publishes a source-segment citation rate only after a segment clears at least 10 observations across at least seven distinct run dates. Below that line, the segment is marked collecting rather than scored. (Machine Relations Index public data)
That design choice is easy to miss. A tool with no evidence floor can feel more complete because every prompt produces a number. But completeness can be cosmetic if the score is based on one run, one model, one geography, or one volatile news day.
Statistical guidance outside AI visibility points in the same direction. NIST's handbook treats proportion intervals as a method attached to assumptions and sampling conditions, not as a decorative number after the fact. AAPOR's survey best-practice guidance similarly emphasizes design, data collection, analysis, and transparency before readers judge a result by size alone. (NIST) (AAPOR)
Source roles change AI visibility rankings
A source-citation score changes when the tool separates owned pages, earned media, communities, social platforms, and reference sites. Two brands can have the same mention count while relying on very different evidence layers. One may be cited through its own blog. Another may be cited through reviews, media coverage, documentation, Reddit threads, analyst pages, or partner sites.
That distinction matters because AI systems synthesize from multiple sources. The original Generative Engine Optimization paper formalized generative engines as systems that gather and summarize information from multiple sources, creating a visibility problem for website and content creators. (GEO paper, arXiv)
For operators, source role is diagnostic. If a brand's owned pages appear but independent sources do not, the correction layer is not just more owned content. It may require earned authority, third-party corroboration, clearer documentation, or better citation architecture.
Freshness windows change AI visibility volatility
A score measured today can disagree with a score measured last week because answer systems, indexes, news context, and retrieval caches change. The September 15 MRI manifest identifies a public artifact generated on September 15, with data observed from May 10 through September 15 and 122 observed days. That release identity keeps the number tied to a specific window. (MRI release manifest)
Freshness is not a footnote. A model or retrieval surface can change behavior after a product launch, a press cycle, a ranking update, or a citation-source update. Without release identity, a team cannot tell whether a score changed because the brand improved, the market moved, the tool changed method, or the prompt sample changed.
The buyer question should be simple: can the vendor reproduce the same report with the same date, engine roster, prompt set, and denominator?
A practical checklist for comparing AI visibility tools
CMOs should compare AI visibility tools by method before comparing scorecards by number. The score is the output. The method is the product.
| Method field | What to ask | Why it changes the score |
|---|---|---|
| Engine roster | Which models and answer surfaces are measured? | Engines cite and summarize differently. |
| Prompt basket | Are prompts brand, category, competitor, or buyer-intent prompts? | Different prompts answer different business questions. |
| Denominator | Is the score based on answer runs, citations, mentions, positions, or traffic? | The denominator defines what the percentage means. |
| Evidence floor | How many runs and dates are required before a score is published? | Thin signal can look precise without replication. |
| Source roles | Are owned, earned, community, reference, and social sources separated? | Correction work depends on the source layer. |
| Freshness window | What dates does the report cover, and can the result be reproduced? | AI answer surfaces change over time. |
| Missingness | Are failed runs, empty answers, and collecting segments disclosed? | Missing data can inflate certainty. |
This checklist does not require every vendor to use the Machine Relations Index method. It requires every vendor to disclose enough method for a buyer to interpret the output.
What brands should do after the score changes
The right response to a changed AI visibility score is diagnosis before correction. If the score fell because a key answer engine stopped citing an earned-media source, the response is different from a decline caused by stale owned pages or a prompt-basket change.
Use a three-layer operating model:
| Layer | Question | Output |
|---|---|---|
| Detection | Are we mentioned, cited, omitted, or misdescribed? | Scorecards, prompt logs, citation lists, sentiment flags |
| Diagnosis | Which engine, prompt, source role, and date changed? | Root-cause map by answer surface and source layer |
| Correction | What public evidence should change? | Updated pages, stronger third-party proof, clearer claims, schema, earned media |
This is where Machine Relations becomes a useful operating frame. Machine Relations is the discipline of making brands legible, retrievable, and credible inside AI-mediated discovery systems. It treats measurement as one layer of a larger source architecture, not as a substitute for the evidence machines need to cite.
FAQ
Why do AI visibility tools show different scores for the same brand?
AI visibility tools show different scores because they often measure different engines, prompts, source roles, dates, and denominators. A brand can be visible in ChatGPT, absent from Google AI Mode, cited through earned media in Perplexity, and merely mentioned without citation in another surface.
Is one AI visibility score the true number?
No single AI visibility score is the true number unless the business question and method are specified. A citation-rate score answers how often sources were cited in observed answer runs; a brand-mention score answers whether the brand appeared in a prompt sample.
What is a good evidence floor for AI visibility measurement?
A good evidence floor discloses how many observations and run dates stand behind the score. MRI v2 uses at least 10 observations across at least seven distinct run dates before publishing a source-segment citation rate, while thinner segments are marked collecting.
Should brands buy an AI visibility tool before fixing source architecture?
Brands should use measurement to guide source architecture, not replace it. A tool can detect omissions, citations, and sentiment, but the correction work usually requires clearer owned pages, better third-party corroboration, consistent entity signals, and more extractable proof.
How should CMOs evaluate an AI visibility vendor?
CMOs should ask for the engine roster, prompt sample, denominator, evidence floor, source-role taxonomy, release date, missing-data policy, and reproducibility rules. If the vendor cannot explain those fields, the headline score is not ready for budget decisions.
Sources
- Machine Relations Research. "Machine Relations Index public data." September 15, 2026. https://machinerelations.ai/data/machine-relations-index.json
- Machine Relations Research. "MRI release manifest." September 15, 2026. https://machinerelations.ai/data/mri-release-manifest.json
- Google Search Central. "Guide to Optimizing for Generative AI Features on Google Search." https://developers.google.com/search/docs/fundamentals/ai-optimization-guide
- Aggarwal et al. "GEO: Generative Engine Optimization." arXiv. https://arxiv.org/abs/2311.09735
- NIST/SEMATECH. "7.2.4.1. Confidence intervals." e-Handbook of Statistical Methods. https://www.itl.nist.gov/div898/handbook/prc/section2/prc241.htm
- AAPOR. "Best Practices for Survey Research." https://aapor.org/standards-and-ethics/best-practices/