TL;DR

AI visibility measurement tracks whether an assistant mentions your brand, links to your website, and describes your business accurately. A useful report keeps those observations separate from visits, leads, and sales. It also preserves the prompts and answers behind every percentage.

You can start with a spreadsheet. The difficult part is deciding what to measure and keeping the method consistent. This guide gives a small marketing team a repeatable baseline without pretending that a few prompts represent every customer.

Define the question the report must answer

“Are we visible in AI?” is too broad. A more useful question is: “When a US buyer compares project-management tools for a ten-person agency, which products appear, which sources support the answer, and what is said about ours?”

Write down the market, language, product category, buyer situation, and platforms in scope. Decide whether you are studying a search-enabled consumer experience or an API response. Keep those datasets separate; they need not produce the same answers.

If the terminology is new, read our GEO, AEO, and SEO explainer first.

Build a small prompt set from real questions

Start with questions collected from sales calls, customer interviews, support tickets, and search data you already have. Group them by the decision they support.

Prompt groupIllustrative questionWhat it helps you examine
ProblemHow should a small agency manage client approvals?Whether your category is introduced
ShortlistWhich project-management tools suit a ten-person agency?Which brands are considered
ComparisonHow do product A and product B handle guest access?Whether the tradeoffs are accurate
VerificationDoes product A support approval history?Whether specific facts can be found
BuyingWhat should I check before choosing a plan?Which requirements shape the decision

These are examples, not evidence of search volume. Do not present a generated prompt list as measured consumer demand. Review the wording with people who know the customers, and remove prompts that exist only to make your brand appear.

Keep branded and unbranded prompts in separate groups. Asking about your company by name is a useful accuracy test, but a poor measure of whether an unfamiliar buyer would discover it.

Record the conditions of every run

Use a fresh conversation for each independent test. Record the platform, visible model or mode, language, location where relevant, date, and whether search was enabled. Note personalization or account conditions you can observe.

Save the exact prompt, full response, and citation URLs. A cropped screenshot of a brand mention leaves too much out. If the tool fails, record the failure rather than silently removing it from the results.

For a manageable starting routine, select 20 prompts and repeat each three times per platform during a defined measurement window. That is an example of a small operational baseline, not a statistically representative market survey. Increase the sample when the decision warrants it, and respect each platform's usage terms.

Keep the prompt set fixed for the comparison period. If you add questions, version the set and report the additions separately.

Calculate metrics with explicit denominators

A brand mention, a recommendation, and a citation are different observations. Define the coding rules before counting.

MetricCalculationWhat it does not establish
Brand mention rateValid responses mentioning the brand / valid responses sampledTotal market awareness
Owned-site citation rateValid responses linking to your domain / valid responses sampledVisits or conversions
Recommendation rateValid responses that positively shortlist the brand / valid responses sampledA stable ranking
Factual error rateReviewed brand-containing responses with a material error / reviewed brand-containing responsesAccuracy across every possible question
Share of tracked brand mentionsYour brand-response mentions / all brand-response mentions in the defined competitor setShare of all AI users or impressions

Count a brand once per response when using these definitions. If a response repeats the same brand five times, that does not create five separate discoveries. Document how ambiguous names, subsidiaries, and product names are handled.

Here is a worked example using invented data. Across 60 valid responses, your brand appears in 18 and your website is cited in 9. Mention rate is 30%; owned-site citation rate is 15%. If the tracked competitor set has 90 brand-response mentions in total, your share is 20%. These percentages describe the sample, not the whole market.

Show the counts beside the rates. Moving from two mentions to four may look dramatic as a percentage change, but it is still a small number of observations.

Add platform data and website outcomes

Microsoft's AI Performance report in Bing Webmaster Tools provides citation activity for supported AI experiences. Its citation counts are not rankings or proof of a visit. Use the report as another source of evidence, with its own coverage and definitions.

Google includes traffic from AI features within its overall Web reporting in Search Console, as explained in its AI features documentation. Avoid presenting ordinary Search Console totals as an isolated AI Overview report.

In website analytics, group identifiable AI referrals and inspect the landing pages and conversions. Keep paid ChatGPT campaigns separate using a consistent tagging convention. Missing referral information means observed AI traffic may be incomplete; it does not justify assigning every unattributed conversion to AI.

Turn observations into a short work list

Read the actual answers before recommending another article. If the assistant cites an old pricing page, correct the source. If it confuses two products, clarify their names and relationships. If the cited comparison explains a requirement your site omits, consider adding that information.

Use a report with four lines per issue:

  1. Observation: What appeared, with a saved response and source.
  2. Interpretation: The most plausible explanation, labeled as a hypothesis.
  3. Action: The specific page or public fact to improve, with an owner.
  4. Review: When and how the same questions will be tested again.

The GEO content audit checklist helps turn that list into page-level work. If the volume becomes difficult to manage, evaluate AI visibility tools against this method rather than buying on the strength of a dashboard score.

Report change without claiming causation

Compare the same prompt cohort, platform, and conditions across periods. Mark model changes, major launches, and edits to the competitor set. Keep new prompts out of the historical comparison until they have their own baseline.

A useful conclusion is specific: “Our domain was cited in 9 of 60 sampled responses this period, compared with 5 of 60 last period; the new comparison page accounted for three citations.” That tells the team what was observed. It does not claim the page caused all of the increase or that the result will persist.

A good AI visibility report ends with a decision: correct a fact, improve a page, investigate a source, or keep observing. More charts are only useful when they help make that decision.

Sources and further reading

AI Performance report in Bing Webmaster ToolsAI features documentation

Put the research to work

Build a practical AI visibility plan for your brand.

Start a conversation