A useful measurement system separates collection quality, answer presence, competitive position, source adoption, and downstream business outcomes. Every rate needs a clear numerator, denominator, and exclusion rule.
Quick answer
Start with these eight metrics:
- Response success rate
- Strict brand mention rate
- Recommendation or preference rate
- Description accuracy
- Official citation rate
- External source adoption
- Answer stability
- AI referral and conversion
Do not count a failed collection as zero visibility. Do not add citations, mentions, traffic, and revenue into a single score unless the weighting has a defensible decision purpose.
UnderAI GEO Workspace exposes the operational metrics and evidence: Coverage, Coverage Change, Visibility, TOP3, Sentiment, Answer Snapshots, brand exposure, and Citations under a defined project, date, platform, and Prompt Group scope. An AI Visibility Audit reviews how those fields should be interpreted for a real brand and traces decisions back to the underlying responses and URLs.
In UnderAI’s current product vocabulary, Sentiment is a per-snapshot Positive / Neutral / Negative state. It replaces the older aggregate-rate concept and should be reported as a state, not a percentage.
1. Response success rate
Formula
Response success rate = successful evaluable response units / planned response units
This is a data-quality metric, not a brand-performance metric. It tells you whether the platform and collection route returned usable evidence.
Report platform failures, empty results, parser failures, and excluded units separately. Visibility rates should normally use successful, evaluable response units as their denominator.
2. Strict brand mention rate
Formula
Strict mention rate = successful units with the exact target brand / successful evaluable units
Resolve aliases and namesakes before calculation. A partial token, unrelated entity, or mention inside a source title should not automatically count as a brand mention in the answer.
What it answers: Is the brand present?
What it cannot answer: Was the brand recommended, described accurately, or supported by evidence?
3. Recommendation and preference rate
Only calculate this for questions that invite a recommendation, comparison, or shortlist.
Classify the brand’s role before aggregating:
- first choice;
- conditional leader;
- peer candidate;
- mentioned but not recommended;
- explicitly not recommended;
- absent.
Example formula
Positive recommendation rate = qualifying successful units where the brand is a first choice, conditional leader, or peer candidate / qualifying successful units
Publish the category distribution with the rate. A single percentage can hide whether the brand is consistently first or merely listed at the end.
4. Description accuracy
Use brand-fact questions and a frozen fact ledger. Score only facts that can be evaluated against current official evidence.
Formula
Description accuracy = correct evaluated facts / all evaluated facts
Track serious inaccuracies separately from omissions. A missing secondary detail and a false claim about the product category should not carry the same risk.
5. Official citation rate
Formula
Official citation rate = successful units citing the official domain / successful evaluable units
Canonicalize URLs and define which domains are official. A citation to a social profile, app store, documentation subdomain, or partner page may need a separate publisher type.
What it answers: How often does the answer use a brand-controlled source?
What it cannot answer: Did a specific page module cause the citation, or did the user visit the site?
6. External source adoption
External citations explain which publishers, communities, analysts, reviews, and expert pages supply evidence around a topic.
Useful views include:
- cited response units by domain;
- unique canonical URLs;
- prompt coverage by source;
- platform coverage by source;
- repeat adoption of the same page across independent responses;
- source type and author or account identity.
Domain frequency does not prove that a site accepts contributions or that publishing there will cause a citation. Treat publication access and citation demand as separate facts.
7. Answer stability
AI answers vary. Repeated samples help distinguish a stable pattern from one result.
Possible measures include:
- mention consistency across repetitions;
- overlap of recommended brands;
- citation URL overlap;
- description fact agreement;
- position or role changes.
When a platform changes its model, interface, or collection route, annotate the series. A method change can look like a performance change.
8. AI referral and conversion
AI referral traffic is a downstream behavior metric. Keep it separate from answer visibility.
Measure:
- identifiable AI referral sessions;
- landing pages;
- engaged sessions;
- qualified leads;
- assisted pipeline or revenue where attribution is defensible.
An answer can influence a decision without sending a click, and a referral can occur without a visible citation in your monitoring sample. Report both without forcing a false one-to-one relationship.
Optional competitive metrics
Share of voice
Define the unit before using the term. It may mean share of qualifying responses with a mention, share of all brand appearances, or share of first-choice recommendations. These are different metrics.
Citation share
This can describe a brand’s official citations as a share of official competitor citations within the same prompt and platform scope. It is not comparable across different panels.
Source diversity
Count unique domains or publisher types only when diversity supports a decision. More domains are not automatically better; a small set of authoritative sources may be more useful than many weak ones.
A minimum reporting table
| Field | Why it must be included |
|---|---|
| Panel version | Prevents different prompt sets from being compared as one trend |
| Platform and route | Separates collection environments |
| Planned / successful / failed units | Exposes data completeness |
| Prompt purpose | Keeps recommendation, brand fact, and research metrics separate |
| Numerator / denominator | Makes every rate reproducible |
| Baseline and comparison date | Shows when change was measured |
| Page or source changes | Supports cautious attribution |
| Raw evidence path | Allows independent review |
What is a good AI visibility score?
There is no universal threshold. A good result depends on question type, competitive set, platform, brand maturity, and business objective.
Use three comparisons instead:
- the same brand against its frozen baseline;
- the brand against relevant competitors within the same response units;
- the result against the intended page or campaign objective.
If a score cannot be traced to the underlying responses, use it as a directional signal, not an acceptance criterion.
Build metrics into action
- Low mention with correct descriptions when mentioned: improve discovery and category evidence.
- High mention with low recommendation: strengthen differentiation and decision proof.
- Good recommendation with low official citation: improve the official owner page and source clarity.
- High external citation concentration: inspect which third-party pages supply the answer and whether the brand is represented accurately.
- High volatility: increase repetitions before changing strategy.
What an UnderAI AI Visibility Audit delivers
The audit turns the metric dictionary into a decision baseline. The client receives:
- the reviewed Prompt Library and fixed measurement scope;
- planned, successful, failed, and excluded Response counts;
- strict Mention, recommendation role, description accuracy, official Citation, and external-source views;
- competitor and platform breakdowns using compatible denominators;
- a map from each priority gap to an official page, external source, or product question;
- the baseline version and a Monitoring & Retesting plan.
The purpose is to identify what should change next, not to maximize a composite score.
Learn how to build the monitoring panel behind these metrics or request an evidence-led AI Visibility Audit.
Worked calculation from an UnderAI baseline
The July 23, 2026 Panel provides a concrete example of why metric denominators must be published.
| Metric | Calculation | Result | Interpretation |
|---|---|---|---|
| Response success rate | 141 successful / 240 planned | 58.8% | Collection completeness, not brand visibility |
| Strict UnderAI mention rate | 0 mentioned / 141 successful | 0.0% | Valid baseline non-mention under a strict entity rule |
| Any structured Citation rate | 93 with Citation / 141 successful | 66.0% | Platform answers often used sources, but not UnderAI |
| UnderAI official Citation rate | 0 official citations / 141 successful | 0.0% | Official-source gap in the measured sample |
| Unique Citation URLs | Deduplicated canonical URLs | 231 | Breadth of source evidence, not source quality |
The 99 failed or missing units do not enter the mention or Citation denominators. No arithmetic can turn an unavailable Google AI Mode result into a measured zero.
This example also shows why a composite “AI visibility score” can mislead. A 66% general Citation rate sounds active, while the brand itself remained absent. Teams need the component metrics before they can choose an action.
Data source and interpretation boundary
The example uses a frozen 16-Prompt, five-route, three-run US English baseline collected on July 23, 2026. Read the AI Citation Rate Benchmarks for platform cuts and the Most Cited Domains study for source distribution. These figures are a category baseline, not a universal target for every brand. Google’s AI feature guidance is the source for the Search Console and indexing boundaries used in this metric system.
Frequently asked questions
Is mention rate the same as visibility?
Mention rate is one visibility measure. It does not capture recommendation position, accuracy, citations, or stability.
Should I combine all platforms into one rate?
Report each platform first. A combined rate is only useful when the platform mix, successful response counts, and weighting are explicit.
Should a platform failure count as zero?
No. It should reduce the response success rate and remain excluded from the brand-visibility denominator.
Is citation rate more important than mention rate?
They answer different questions. Mention rate measures presence. Citation rate measures source adoption. The priority depends on whether the current problem is discovery, trust, accuracy, or traffic.
Can I use AI referral traffic as proof that AEO worked?
Referral growth is valuable evidence, but it does not identify the cause by itself. Review landing pages, timing, answer changes, other campaigns, and attribution limits before making a causal claim.
