The short answer
An AI visibility score summarizes how often and how prominently a brand appeared across a defined set of questions, providers, markets, and collection dates. It is useful only when you can inspect that scope, the completed-answer denominator, the underlying responses, and the scoring rules. It is evidence about a measured sample—not a universal percentage of every AI answer or every potential customer.
What businesses notice
Common signs of the problem
- A dashboard shows one percentage but does not identify the questions or providers behind it.
- Failed provider responses are counted as brand omissions or silently removed.
- A higher score looks encouraging, but nobody can explain which buyer questions changed.
- Two vendors report different scores for the same brand and both call the result definitive.
01
Start with the measurement contract
Before interpreting the number, define exactly what was observed. A useful score is the final line of a measurement contract, not a substitute for one.
Name the question set
List the exact buyer questions and their versions. A score built from category-discovery prompts measures something different from one built from comparisons, local recommendations, or troubleshooting questions.
Name the provider and market scope
Record which AI providers were tested, when they were tested, and any language, location, or audience context. ChatGPT can use search and location context, while other products may retrieve and present sources differently.
Review ChatGPT Search behavior ↗Define the counted outcome
Say whether the score represents any mention, a recommendation, first position, citation coverage, factual accuracy, or a weighted combination. These outcomes answer different business questions and should not be blended without visible rules.
Freeze the denominator
Show planned observations, completed answers, failures, and exclusions. Calculate brand performance from completed, eligible answers and report collection coverage beside it so infrastructure failures never become fake visibility losses.
02
Understand the common score types
No single formula is mandatory. The right calculation depends on the decision, but the name and formula should make the measured behavior obvious.
Mention coverage
Completed answers that mention the tracked brand divided by completed eligible answers. This is the clearest broad presence measure, but it does not tell you whether the brand was recommended, criticized, or merely referenced.
Recommendation coverage
Completed answers that present the brand as a suitable option divided by completed eligible answers. The evaluation rubric must distinguish an actual recommendation from a neutral mention or a warning.
Position-weighted visibility
A transparent weighting can distinguish first choice from later inclusion. Publish the weights and also show raw counts; otherwise a small formula change can create an apparent trend with no underlying answer change.
Citation and accuracy measures
Citation coverage asks whether owned or relevant evidence appears. Accuracy asks whether material claims match a source of truth. Keep both separate from visibility because a visible brand can still be unsupported or described incorrectly.
See the brand-perception audit →03
Work through a transparent example
Suppose a company tests 10 frozen buyer questions in ChatGPT and Gemini. That creates 20 planned question-provider observations.
Report collection coverage first
If 18 answers complete and two fail, collection coverage is 18/20, or 90%. The performance denominator is 18 completed answers. The failures remain visible but are not counted as brand omissions.
Calculate the observable rates
If the brand appears in 9 completed answers, mention coverage is 9/18, or 50%. If it is genuinely recommended in 5, recommendation coverage is 5/18, or 28% when rounded to a whole percent.
Keep the breakdown actionable
A 50% mention rate becomes useful when the team sees 4/4 problem prompts, 3/6 category prompts, and 2/8 comparison prompts. The comparison gap is a clearer work queue than the blended score.
Describe only the tested sample
The defensible conclusion is that the brand appeared in half of these 18 completed observations. It does not mean the brand owns half of all AI visibility. NIST distinguishes performance on a fixed benchmark from broader generalization and recommends explicit measurement targets and uncertainty.
Read NIST AI evaluation guidance ↗04
Interrogate a score before trusting it
A score can be mathematically correct and still be unhelpful if its scope, sampling, or labels do not match the decision you need to make.
Can you inspect the evidence?
Every material result should lead back to the exact question, provider output, timestamp, citations or source links when available, and the structured evaluation. OpenAI notes that searched responses may include inline citations or a Sources panel; preserve what the provider actually returned.
Are the prompts representative?
A hand-picked set of easy brand-name prompts can inflate visibility. A set built only from adversarial comparisons can depress it. Use real buyer jobs, label exploratory questions, and disclose how the portfolio was chosen.
Are different providers being flattened?
A blended total can hide that the brand is consistently visible in one provider and absent in another. Show provider-level results before the rollup so ecosystem-specific gaps remain visible.
Compare provider disagreement →Did the method change?
Prompt edits, provider additions, new markets, rubric changes, and revised weighting can all move the number. Version the methodology and mark breaks in the series instead of presenting unlike measurements as a continuous trend.
05
Use the score as a doorway, not a verdict
The value of the number is that it directs attention to specific evidence. The score itself does not improve visibility.
Open the largest material gap
Move from the rollup to the question group with repeated omissions, competitor displacement, inaccurate claims, or missing owned evidence. Inspect the complete answers before proposing work.
Form a testable hypothesis
Connect the observed gap to a plausible evidence issue: unclear category language, missing service detail, weak comparison proof, inconsistent location facts, or inaccessible source material. Record alternative explanations too.
Change one meaningful thing
Improve the relevant public evidence, then repeat the same measurement contract. Do not rewrite the prompt portfolio after seeing the baseline merely to produce a more favorable score.
Begin with one inspectable question
Use the free AI visibility snapshot to compare one buyer-style question across ChatGPT and Gemini. Read both full answers, note competitors and evidence, and decide whether a larger baseline would change a real business decision.
Run a free two-provider snapshot →06
Keep the score honest over time
A trend earns trust when readers can separate brand movement from collection noise, provider change, and methodology change.
Retain counts beside percentages
“Nine of 18 completed answers” is more informative than “50%.” Counts expose small samples and make changes in the denominator obvious.
Separate Google-owned visibility data
Google documents impressions, pages, countries, devices, and dates for eligible generative AI performance reporting in Search Console. Treat that first-party surface as complementary evidence rather than mixing it silently with answer-level provider tests.
Review Google’s generative AI reports ↗Annotate operational events
Record provider failures, site releases, major content changes, market events, and collection changes on the same timeline. A score movement without context is an observation, not an explanation.
Escalate only repeatable findings
Treat a one-period shift as a review signal. Confirm it across repeated collections, related questions, or additional evidence before making a costly content or positioning decision.
Choose a monitoring cadence →Continue the research
Related AI visibility guides
Primary sources
Primary documentation used for this guide
AI products, search behavior, and platform policies change. Check these maintained first-party sources before making technical decisions.
- ChatGPT Search ↗
OpenAI Help Center
First-party guidance on web search, citations, source panels, location context, crawl eligibility, and the absence of guaranteed top placement.
- Expanding the AI Evaluation Toolbox with Statistical Models ↗
National Institute of Standards and Technology
Primary evaluation guidance distinguishing fixed-benchmark results from broader generalization and emphasizing explicit measurement targets and uncertainty.
- Introducing Search Generative AI performance reports in Search Console ↗
Google Search Central
First-party description of dedicated generative AI visibility reporting, including impressions, pages, countries, devices, dates, and rollout limitations.
- AI features and your website ↗
Google Search Central
Maintained guidance on inclusion in Google AI features and measuring site performance through Search Console and analytics.
Measure your visibility
Turn the questions in this guide into an evidence-backed baseline.
Get AI Visibility ScoreTalk to us
Have a visibility question?
Tell us what your team is trying to measure or improve.
Contact AI Brand Lens →Keep learning
Explore every guide.
Browse practical answers about AI visibility, competitors, and measurement.
View all guides →