AI visibility question design

How many prompts do I need to track AI visibility?

Choose a question set that covers real buyer decisions, then test whether adding more questions changes what you learn. There is no defensible universal prompt count.

10 min read Practical guidePublished September 19, 2026

The short answer

Track enough distinct buyer questions to cover the decisions, audiences, markets, and providers you actually care about. Start with a small, deliberately varied pilot; keep a stable core for comparisons; and expand where new questions reveal new competitors, omissions, inaccurate claims, or source patterns. Ten near-identical prompts can be less useful than six that represent different buyer jobs. Report the exact tested set and completed-answer denominator, never a claim that it represents all possible AI searches.

What businesses notice

Common signs of the problem

  • The dashboard looks strong because nearly every prompt includes your brand name.
  • A team adds synonyms until the count feels impressive but learns nothing new.
  • One high-value buyer segment or market is absent from the question set.
  • A score changes after prompts are replaced, but the report presents it as brand movement.

01

Decide what the question set must represent

The unit is a buyer question in a stated context, not a keyword. Write down the business decision first: discovery, positioning, competitor analysis, claim accuracy, or change after a website update.

Define the eligible buyer journey

List the tasks people plausibly ask an assistant to help with: discovering a category, comparing options, checking fit, verifying a claim, or choosing a provider. Exclude questions where your offering is not legitimately relevant.

Name the audience and market

A procurement lead, end user, and technical evaluator may ask materially different questions. Location, language, product tier, and service area can also change the set. Include a dimension only when it could alter a real decision; do not multiply every combination mechanically.

See why location needs its own scope →

Choose the observed outcome

Decide whether you need to measure mention, recommendation, shortlist inclusion, competitor framing, citations, or factual accuracy. The question portfolio should expose that outcome. A brand-name fact check cannot stand in for an unaided category recommendation.

Keep the target of inference honest

A fixed set can describe performance on those exact questions. It does not automatically estimate performance across every possible customer question. NIST distinguishes fixed-benchmark measurement from generalization to a wider population; the latter needs explicit sampling assumptions.

Read the NIST evaluation paper ↗

02

Build coverage before you choose a number

Use a simple coverage ledger with columns for buyer task, decision stage, audience, market, business priority, exact wording, and why the question belongs. A row without a reason is a candidate for removal.

Seed questions from real language

Start with sales calls, support conversations, site-search terms, customer interviews, and the objections your team hears. Rewrite them as natural requests a person could actually ask. An invented “best brand” prompt chosen only because it flatters your company is a poor sampling unit.

Cover materially different jobs

A useful starter mix might include category discovery, a concrete problem, fit for a use case, a comparison, and a verification question. Those are design categories, not a required quota. OpenAI’s evaluation guidance likewise calls for diverse scenarios and a mix of real-world and expert-created cases.

Review OpenAI evaluation guidance ↗

Separate aided from unaided questions

“Is Acme suitable for my team?” tests description and fit after the buyer names Acme. “Which tools should my team consider?” tests whether Acme enters the candidate set at all. Both can matter, but combining them without labels can make visibility look artificially high.

Avoid synonym inflation

Changing “best” to “top” or rearranging an adjective may be useful in a sensitivity study, but it does not automatically cover a new buyer need. Keep wording variants in a separate experiment unless they represent meaningfully different intent.

03

Run an illustrative 12-question pilot

This is an example design, not a minimum or an AIBL customer result. Suppose a B2B team wants to understand discovery and evaluation for one product in one market.

Allocate by decision, not symmetry

Choose four category or problem questions, four use-case or role questions, and four comparison or verification questions. Write the buyer rationale beside each. If the product has two truly different buyer groups, split the twelve deliberately rather than giving every segment one token question.

Make the observation count visible

Twelve questions across ChatGPT and Gemini create 24 planned question-provider cells per run. Two repeats create 48 planned observations, not 48 independent buyer intents. Record completed, failed, and excluded cells separately; never score a provider failure as an omission.

Inspect the answer, not just a percentage

For each completed cell, capture the prompt, provider, time, full answer, brand inclusion, recommendation role, competitors, citations when shown, and material factual errors. OpenAI warns that searched answers and citations can be incomplete or wrong, so cited claims still need source verification.

Read ChatGPT Search guidance ↗

Write a bounded conclusion

“Our brand was recommended in 7 of 22 completed cells on this pilot set” is an inspectable result. “We own 32% of AI search” is not supported by the same evidence. Break the result down by buyer job before deciding where to investigate.

04

Expand with a saturation check

The right next prompt is the one most likely to change a decision. Add questions in small, predeclared batches rather than filling an arbitrary target number.

Look for a new finding type

After the pilot, add a question from an under-covered audience, stage, use case, or market. Ask whether it reveals a new omission, competitor, inaccurate claim, rationale, citation pattern, or provider split. If several well-chosen additions produce no new decision-relevant pattern, the current portfolio may be sufficient for this decision.

Do not mistake stability for completeness

A plateau within one category says little about an untested market or buyer role. Saturation is conditional on the strata you defined. Document untested segments as blind spots rather than claiming comprehensive coverage.

Use repeats for a different question

Repeating the same prompt helps assess whether an observed answer is stable. Adding a distinct prompt expands intent coverage. Both consume collection capacity, but they answer different questions and should be reported separately.

Choose a monitoring cadence →

Stop when marginal insight is low

If a new batch only restates existing patterns and cannot change a content, product, sales, or measurement decision, stop expanding for now. Revisit the set when the business adds a market, offering, buyer role, or competitive question.

05

Protect the baseline as the set evolves

Use a stable core for trend comparisons and an exploratory set for learning. A prompt can graduate into the core, but the reporting break must be explicit.

Freeze a core with owners

Keep the exact wording, audience, market, eligibility rule, and outcome rubric for questions used in the headline trend. Assign an owner to review whether each still matches a real buyer decision.

Version every meaningful change

A new budget, location, product tier, named competitor, or requested format can change what is being tested. Save the previous wording and start a new version. Do not splice unlike samples into one continuous score.

Interpret a score with its denominator →

Review selection bias

Before reporting, scan for clusters of brand-name prompts, easy questions, or variants created after seeing favorable results. Also look for important buyer objections that the set omits. Record why each retained question matters.

Begin with one honest observation

A free AI visibility snapshot can show the full ChatGPT and Gemini answers to one buyer-style question. Use it as a first evidence check, then build a broader set only if the result exposes a decision worth tracking.

Run the free AI visibility snapshot →

Continue the research

Related AI visibility guides

Primary sources

Primary documentation used for this guide

AI products, search behavior, and platform policies change. Check these maintained first-party sources before making technical decisions.

  • Evaluation best practices

    OpenAI Platform Documentation

    First-party advice on defining evaluation objectives, collecting diverse test cases, identifying input variation, and expanding evaluation sets over time.

  • Expanding the AI Evaluation Toolbox with Statistical Models

    National Institute of Standards and Technology

    Primary research distinguishing measurement on a fixed benchmark from generalization to a wider item population and emphasizing explicit assumptions.

  • Searching the web with ChatGPT

    OpenAI Help Center

    First-party guidance on searched answers, citations, query rewriting, location context, and the need to verify source claims.

Measure your visibility

Turn the questions in this guide into an evidence-backed baseline.

Get AI Visibility Score

Talk to us

Have a visibility question?

Tell us what your team is trying to measure or improve.

Contact AI Brand Lens →

Keep learning

Explore every guide.

Browse practical answers about AI visibility, competitors, and measurement.

View all guides →