Agency workflow guide

How should agencies report AI visibility to clients?

A client-ready reporting method that keeps every AI visibility metric connected to the questions, provider answers, citations, limits, and decisions behind it.

15 min read Practical guidePublished August 5, 2026

The short answer

Report AI visibility as a scoped set of observations, not a universal rank. Agree on the client, market, buyer questions, providers, and schedule before collection; preserve every completed answer and failure; calculate metrics with visible denominators; and connect each material finding to its evidence and next decision. A useful agency report tells the client what was tested, what changed, what did not, how certain the team is, and what to investigate next—without promising placement or attributing movement to an unproven cause.

What businesses notice

Common signs of the problem

  • A monthly deck shows one visibility score but not the questions or answers behind it.
  • The prompt set changes whenever the agency wants a fresher or more favorable story.
  • Provider failures, citations, and contradictory results disappear from the client summary.
  • A before-and-after chart implies that one site edit caused an AI recommendation change.

01

Start with a reporting contract

Before collecting answers, write a one-page measurement contract the client can understand. This prevents the report from expanding or changing after the results are known.

Name the decision the report supports

Define whether the client needs a discovery baseline, competitor review, perception check, citation audit, market comparison, or post-change observation. One report can contain several measures, but it should have one primary decision so every chart earns its place.

Decide whether monitoring is material →

Freeze the scope before scanning

Record the tracked brand and aliases, relevant competitors, market or service area, buyer audience, question set, providers, collection window, and completion rules. If the scope changes later, version it and avoid presenting the new set as a like-for-like trend.

Define every reported metric

Specify what counts as a mention, recommendation, first position, citation, owned citation, competitor, accurate claim, and completed response. Also state the denominator. “Recommended in 6 of 14 completed answers” is reviewable; “43% AI visibility” is not unless the report explains how it was calculated.

Review the evidence-first methodology →

Set the limits in plain language

Explain that generated answers can vary by question, provider, model, retrieved sources, location, personalization, and time. OpenAI says ChatGPT Search can rewrite a prompt into targeted searches and may use location context; neither an agency nor a measurement platform can promise a permanent placement.

Read the current ChatGPT Search guidance ↗

02

Build a client question portfolio

A reporting program should represent how the client is plausibly discovered and evaluated. It should not be a bag of near-duplicate prompts designed to inflate the sample or manufacture a trend.

Organize questions by buyer job

Use a compact mix such as category discovery, problem solving, comparison, qualification, location, and verification. Attach an owner and business rationale to each question. If nobody would act differently based on the answer, remove it from recurring reporting.

Keep client relevance explicit

Only score the brand against questions where it has a legitimate offering, market presence, and audience fit. An omission from an irrelevant category is not a useful failure, and an appearance in a vanity prompt is not evidence of meaningful demand coverage.

Separate stable and exploratory questions

Stable questions support comparison over time. Exploratory questions investigate new services, markets, competitors, or buyer language but stay outside the headline trend until the client intentionally adopts them. Label both groups in the report.

Version meaningful prompt changes

Preserve the exact wording and mark changes to geography, constraints, audience, freshness, or requested format. Small wording edits can change the generated answer. A client should be able to distinguish provider movement from a redesigned test.

See the general tracking method →

03

Preserve an evidence bundle for every answer

The deliverable is not just the slide deck. It is the recoverable record that lets an account lead, specialist, or client reviewer inspect how any conclusion was reached.

Keep the collection context

Store the client or brand, exact question and version, provider and available model context, collection timestamp, market context, completion status, and methodology version. Mark failed and unavailable responses separately from genuine brand omissions.

Retain the complete answer

Preserve the full provider output before extracting mentions, order, competitors, claims, and citations. A screenshot can be useful in a presentation, but it should link back to text and metadata that remain searchable and reviewable.

Capture sources without overstating them

Record links the provider returned and the claim or answer area they appear to support. ChatGPT Search may show inline citations or a Sources panel. Gemini says sources are available only for some responses. Absence of a visible source is a reportable result, not proof that the answer used no information.

Review Gemini source behavior ↗

Keep evaluation separate from evidence

Store structured interpretation—such as mention status, recommendation position, competitor names, and claim accuracy—alongside rather than over the answer. When automation and a human reviewer disagree, preserve the correction and reason instead of silently rewriting history.

04

Use a four-panel client report

A compact report can answer the client’s main questions without collapsing different signals into one score. Lead with decisions and exceptions, then let readers drill into the evidence.

Panel 1: coverage and visibility

Show planned, completed, failed, and excluded observations first. Then report mentions and recommendations by question and provider with denominators. Never count a provider failure as a brand omission, and never hide it to improve the completion rate.

Panel 2: competitive displacement

List which competitors appeared, on which buyer jobs, and in what role or order. Distinguish a competitor that repeatedly replaces the client from one that appeared once in an unordered list. Link material displacement back to the complete answers.

Diagnose competitor recommendations →

Panel 3: sources and claim accuracy

Report citation coverage separately from brand visibility. Identify owned, third-party, competitor, and unknown sources where useful, then flag material statements about services, locations, pricing, qualifications, or policies that require verification against a current source of truth.

Audit AI brand perception →

Panel 4: decisions and next tests

Close with a short queue: finding, evidence, business importance, owner, proposed action, success measure, and next observation date. Separate content or technical work from items that need client fact-checking, legal review, sales context, or no action.

05

Work through a 16-cell example

Suppose an agency monitors one client with eight frozen buyer questions across ChatGPT and Gemini. That creates 16 planned question-provider cells. The collection returns 14 completed answers and two provider failures.

Report completion before performance

Lead with “14 of 16 planned observations completed; two Gemini responses failed.” This tells the client exactly which evidence exists. Do not describe the two failures as omissions or quietly remove their questions from the denominator.

Show the observed rates

If the client is mentioned in 9 completed answers and recommended in 6, report 9/14 mention coverage (64%) and 6/14 recommendation coverage (43%). If it appears first in 3 answers, report 3/14 first-position observations (21%). Keep the fractions beside the percentages.

Break the rollup into buyer jobs

The total becomes actionable only when split. The client might be recommended in all four problem-solving answers, one of four comparison answers, and one of six category-discovery answers. The report should direct attention to comparison and discovery evidence—not celebrate one blended average.

Qualify the conclusion

Say “In this fixed 14-answer sample, the client appeared more often for problem-solving questions than category discovery.” Do not say “AI prefers the client for problems” or generalize the rate to every buyer question. NIST distinguishes performance on a fixed benchmark from performance across a broader question population and recommends making the measurement target and uncertainty explicit.

Review NIST AI evaluation guidance ↗

06

Report change without inventing causality

A recurring report becomes valuable when it explains comparable movement. It becomes misleading when every favorable change is credited to agency work and every unfavorable change is blamed on a provider update.

Compare only like with like

Before calculating change, confirm that the brand scope, questions, markets, providers, completion rules, and metric definitions are comparable. If provider coverage or prompt versions changed, show a break in the series or provide both old-scope and new-scope views.

Maintain a change ledger

Record the baseline observation, evidence checked, working hypothesis, intervention, publication or deployment date, next collection, and observed result. This turns “we optimized for AI” into a traceable sequence and keeps alternative explanations visible.

Use proportional language

Write “recommendation coverage rose from 5/14 to 8/14 after the service-page update” when that is what the record shows. Unless the design supports causal inference, do not claim the edit caused the gain. Repeat the observation before presenting a one-period move as durable.

Preserve unchanged and negative results

A client report should show stable gaps, new errors, lost citations, and competitor gains as clearly as improvements. Negative results can prevent wasted content work or reveal that the initial hypothesis was wrong. They are part of the value, not failed marketing.

07

Review claims before the report travels

Agency reports often become sales slides, executive updates, case studies, and public marketing. Add a review step before a scoped observation turns into a broad promotional claim.

Flag claims that imply more than the test

Phrases such as “#1 in AI,” “the most visible brand,” “AI market leader,” or “our work increased leads” can imply broader provider, question, time, or causality coverage than the evidence supports. Replace them with the exact sample, period, providers, and observable outcome.

Require evidence before objective promotion

The FTC’s advertising substantiation policy says advertisers and ad agencies should have a reasonable basis for objective claims before dissemination. Treat that as a minimum claim-discipline principle, route regulated or high-risk claims to qualified counsel, and do not present this reporting workflow as legal advice.

Read the FTC substantiation policy ↗

Protect client and customer information

Use the minimum data needed for visibility reporting. Avoid placing personal data, confidential strategy, unreleased claims, or customer records into buyer prompts or editorial evidence bundles. Follow the client’s access, retention, approval, and disclosure requirements.

Label examples and screenshots

Mark representative artifacts, redacted examples, and live client evidence accurately. Include collection dates and avoid presenting a favorable answer as current after it has gone stale. Obtain the client permissions needed before using any result in a public case study.

08

Turn reporting into a useful client cadence

The best agency rhythm reserves meeting time for decisions rather than reading every chart aloud. Use the report to make the next month’s evidence plan smaller and sharper.

Triage before the meeting

Have the account lead and specialist review failures, material claim errors, provider disagreements, and large movements before presenting them. Resolve obvious evaluation mistakes and attach the original answers to anything the client may challenge.

Run a three-question review

Ask: What changed enough to matter? What evidence best explains the current hypothesis? What single action or follow-up test deserves ownership? Defer low-impact anomalies unless they repeat across questions, providers, or periods.

Start a client with one real question

Use the free AI visibility snapshot to compare how ChatGPT and Gemini answer one buyer-style question for a prospective or existing client. Review the complete outputs and competitors together before proposing a larger monitoring scope.

Run a free two-provider snapshot →

Scale only when the pilot changes a decision

If the snapshot reveals a material omission, competitor pattern, perception error, or source gap, define the smallest stable question portfolio needed to monitor it. If it reveals nothing actionable, document the result and revisit later instead of manufacturing a recurring deliverable.

Explore AI visibility monitoring →

Continue the research

Related AI visibility guides

Primary sources

Primary documentation used for this guide

AI products, search behavior, and platform policies change. Check these maintained first-party sources before making technical decisions.

  • ChatGPT Search

    OpenAI Help Center

    Current first-party guidance on search-query rewriting, citations and sources, location context, crawl eligibility, and the absence of guaranteed top placement.

  • View related sources from Gemini Apps

    Google Gemini Apps Help

    First-party explanation of how Gemini may show sources and related links, including that sources are not available for every response.

  • Expanding the AI Evaluation Toolbox with Statistical Models

    National Institute of Standards and Technology

    NIST guidance on explicitly defining measurement targets, distinguishing fixed-benchmark results from broader generalization, and communicating uncertainty.

  • Policy Statement Regarding Advertising Substantiation

    Federal Trade Commission

    Primary policy on possessing a reasonable basis before disseminating objective advertising claims, including the responsibility of advertisers and ad agencies.

Measure your visibility

Turn the questions in this guide into an evidence-backed baseline.

Get AI Visibility Score

Talk to us

Have a visibility question?

Tell us what your team is trying to measure or improve.

Contact AI Brand Lens →

Keep learning

Explore every guide.

Browse practical answers about AI visibility, competitors, and measurement.

View all guides →