A useful measurement system separates collection quality, answer presence, competitive position, source adoption, and downstream business outcomes. Every rate needs a clear numerator, denominator, and exclusion rule.

Quick answer

Start with these eight metrics:

  1. Response success rate
  2. Strict brand mention rate
  3. Recommendation or preference rate
  4. Description accuracy
  5. Official citation rate
  6. External source adoption
  7. Answer stability
  8. AI referral and conversion

Do not count a failed collection as zero visibility. Do not add citations, mentions, traffic, and revenue into a single score unless the weighting has a defensible decision purpose.

UnderAI GEO Workspace exposes the operational metrics and evidence: Coverage, Coverage Change, Visibility, TOP3, Sentiment, Answer Snapshots, brand exposure, and Citations under a defined project, date, platform, and Prompt Group scope. An AI Visibility Audit reviews how those fields should be interpreted for a real brand and traces decisions back to the underlying responses and URLs.

In UnderAI’s current product vocabulary, Sentiment is a per-snapshot Positive / Neutral / Negative state. It replaces the older aggregate-rate concept and should be reported as a state, not a percentage.

1. Response success rate

Formula

Response success rate = successful evaluable response units / planned response units

This is a data-quality metric, not a brand-performance metric. It tells you whether the platform and collection route returned usable evidence.

Report platform failures, empty results, parser failures, and excluded units separately. Visibility rates should normally use successful, evaluable response units as their denominator.

2. Strict brand mention rate

Formula

Strict mention rate = successful units with the exact target brand / successful evaluable units

Resolve aliases and namesakes before calculation. A partial token, unrelated entity, or mention inside a source title should not automatically count as a brand mention in the answer.

What it answers: Is the brand present?

What it cannot answer: Was the brand recommended, described accurately, or supported by evidence?

3. Recommendation and preference rate

Only calculate this for questions that invite a recommendation, comparison, or shortlist.

Classify the brand’s role before aggregating:

  • first choice;
  • conditional leader;
  • peer candidate;
  • mentioned but not recommended;
  • explicitly not recommended;
  • absent.

Example formula

Positive recommendation rate = qualifying successful units where the brand is
a first choice, conditional leader, or peer candidate / qualifying successful units

Publish the category distribution with the rate. A single percentage can hide whether the brand is consistently first or merely listed at the end.

4. Description accuracy

Use brand-fact questions and a frozen fact ledger. Score only facts that can be evaluated against current official evidence.

Formula

Description accuracy = correct evaluated facts / all evaluated facts

Track serious inaccuracies separately from omissions. A missing secondary detail and a false claim about the product category should not carry the same risk.

5. Official citation rate

Formula

Official citation rate = successful units citing the official domain / successful evaluable units

Canonicalize URLs and define which domains are official. A citation to a social profile, app store, documentation subdomain, or partner page may need a separate publisher type.

What it answers: How often does the answer use a brand-controlled source?

What it cannot answer: Did a specific page module cause the citation, or did the user visit the site?

6. External source adoption

External citations explain which publishers, communities, analysts, reviews, and expert pages supply evidence around a topic.

Useful views include:

  • cited response units by domain;
  • unique canonical URLs;
  • prompt coverage by source;
  • platform coverage by source;
  • repeat adoption of the same page across independent responses;
  • source type and author or account identity.

Domain frequency does not prove that a site accepts contributions or that publishing there will cause a citation. Treat publication access and citation demand as separate facts.

7. Answer stability

AI answers vary. Repeated samples help distinguish a stable pattern from one result.

Possible measures include:

  • mention consistency across repetitions;
  • overlap of recommended brands;
  • citation URL overlap;
  • description fact agreement;
  • position or role changes.

When a platform changes its model, interface, or collection route, annotate the series. A method change can look like a performance change.

8. AI referral and conversion

AI referral traffic is a downstream behavior metric. Keep it separate from answer visibility.

Measure:

  • identifiable AI referral sessions;
  • landing pages;
  • engaged sessions;
  • qualified leads;
  • assisted pipeline or revenue where attribution is defensible.

An answer can influence a decision without sending a click, and a referral can occur without a visible citation in your monitoring sample. Report both without forcing a false one-to-one relationship.

Optional competitive metrics

Share of voice

Define the unit before using the term. It may mean share of qualifying responses with a mention, share of all brand appearances, or share of first-choice recommendations. These are different metrics.

Citation share

This can describe a brand’s official citations as a share of official competitor citations within the same prompt and platform scope. It is not comparable across different panels.

Source diversity

Count unique domains or publisher types only when diversity supports a decision. More domains are not automatically better; a small set of authoritative sources may be more useful than many weak ones.

A minimum reporting table

FieldWhy it must be included
Panel versionPrevents different prompt sets from being compared as one trend
Platform and routeSeparates collection environments
Planned / successful / failed unitsExposes data completeness
Prompt purposeKeeps recommendation, brand fact, and research metrics separate
Numerator / denominatorMakes every rate reproducible
Baseline and comparison dateShows when change was measured
Page or source changesSupports cautious attribution
Raw evidence pathAllows independent review

What is a good AI visibility score?

There is no universal threshold. A good result depends on question type, competitive set, platform, brand maturity, and business objective.

Use three comparisons instead:

  1. the same brand against its frozen baseline;
  2. the brand against relevant competitors within the same response units;
  3. the result against the intended page or campaign objective.

If a score cannot be traced to the underlying responses, use it as a directional signal, not an acceptance criterion.

Build metrics into action

  • Low mention with correct descriptions when mentioned: improve discovery and category evidence.
  • High mention with low recommendation: strengthen differentiation and decision proof.
  • Good recommendation with low official citation: improve the official owner page and source clarity.
  • High external citation concentration: inspect which third-party pages supply the answer and whether the brand is represented accurately.
  • High volatility: increase repetitions before changing strategy.

What an UnderAI AI Visibility Audit delivers

The audit turns the metric dictionary into a decision baseline. The client receives:

  • the reviewed Prompt Library and fixed measurement scope;
  • planned, successful, failed, and excluded Response counts;
  • strict Mention, recommendation role, description accuracy, official Citation, and external-source views;
  • competitor and platform breakdowns using compatible denominators;
  • a map from each priority gap to an official page, external source, or product question;
  • the baseline version and a Monitoring & Retesting plan.

The purpose is to identify what should change next, not to maximize a composite score.

Learn how to build the monitoring panel behind these metrics or request an evidence-led AI Visibility Audit.

Worked calculation from an UnderAI baseline

The July 23, 2026 Panel provides a concrete example of why metric denominators must be published.

MetricCalculationResultInterpretation
Response success rate141 successful / 240 planned58.8%Collection completeness, not brand visibility
Strict UnderAI mention rate0 mentioned / 141 successful0.0%Valid baseline non-mention under a strict entity rule
Any structured Citation rate93 with Citation / 141 successful66.0%Platform answers often used sources, but not UnderAI
UnderAI official Citation rate0 official citations / 141 successful0.0%Official-source gap in the measured sample
Unique Citation URLsDeduplicated canonical URLs231Breadth of source evidence, not source quality

The 99 failed or missing units do not enter the mention or Citation denominators. No arithmetic can turn an unavailable Google AI Mode result into a measured zero.

This example also shows why a composite “AI visibility score” can mislead. A 66% general Citation rate sounds active, while the brand itself remained absent. Teams need the component metrics before they can choose an action.

Data source and interpretation boundary

The example uses a frozen 16-Prompt, five-route, three-run US English baseline collected on July 23, 2026. Read the AI Citation Rate Benchmarks for platform cuts and the Most Cited Domains study for source distribution. These figures are a category baseline, not a universal target for every brand. Google’s AI feature guidance is the source for the Search Console and indexing boundaries used in this metric system.

Frequently asked questions

Is mention rate the same as visibility?

Mention rate is one visibility measure. It does not capture recommendation position, accuracy, citations, or stability.

Should I combine all platforms into one rate?

Report each platform first. A combined rate is only useful when the platform mix, successful response counts, and weighting are explicit.

Should a platform failure count as zero?

No. It should reduce the response success rate and remain excluded from the brand-visibility denominator.

Is citation rate more important than mention rate?

They answer different questions. Mention rate measures presence. Citation rate measures source adoption. The priority depends on whether the current problem is discovery, trust, accuracy, or traffic.

Can I use AI referral traffic as proof that AEO worked?

Referral growth is valuable evidence, but it does not identify the cause by itself. Review landing pages, timing, answer changes, other campaigns, and attribution limits before making a causal claim.