GEO measurement evaluates how a brand appears in AI-generated answers for a defined prompt set. It should distinguish visibility, citations, competitive presence, factual framing, website outcomes, and business outcomes instead of reducing the program to one score.
The first rule is consistency: define the observation, denominator, competitor set, platform scope, and test period before calculating a rate.
Start With the Measurement Unit
Use one prompt-platform-run observation as the base unit. For example, testing 20 prompts across three AI experiences creates 60 planned observations for one run.
Classify an observation as valid when the platform returns an answer that can be evaluated. Record errors, refusals, unavailable modes, and incomplete captures as excluded observations rather than silently counting them as brand absences.
Some metrics need a narrower eligible denominator. Citation rate, for example, should only use valid observations from experiences where visible citations or source links are available.
1. Mention Rate
Mention rate measures how often the brand appears in valid answers.
Formula:
Mention rate = valid answers with a qualified brand mention ÷ all valid answers × 100
If 18 of 60 valid answers mention the brand, the mention rate is 30%.
Define aliases before reviewing answers. Exclude incidental words that match the brand name but do not refer to the entity.
2. Brand Citation Rate
Brand citation rate measures how often an eligible answer contains at least one visible citation to an approved brand-controlled domain.
Formula:
Brand citation rate = citation-eligible answers with a brand-domain citation ÷ all citation-eligible valid answers × 100
If 7 of 50 citation-eligible answers cite the official domain, the brand citation rate is 14%.
Normalize redirects, subdomains, tracking parameters, and canonical domains before counting. Count an answer once even when it cites several pages on the same approved domain. Page-level citation frequency can be reported separately.
3. Share of Answer
Share of answer compares the tracked brand's qualified appearances with appearances by a fixed set of named competitors.
Formula:
Share of answer = tracked-brand appearances ÷ appearances by the tracked brand and defined competitors × 100
Count each brand at most once per answer unless a different rule is declared in advance. If the tracked brand appears in 20 observations and the competitor set produces 80 total qualified appearances, including those 20, share of answer is 25%.
The competitor set must remain stable between periods. Adding a large competitor later changes the denominator and breaks the comparison unless the historical data is recalculated.
4. Prompt Test Coverage
Prompt test coverage is a data-quality metric. It shows whether the planned matrix was completed.
Formula:
Prompt test coverage = completed prompt-platform-run cells ÷ planned cells × 100
If 114 of 120 planned cells were completed, test coverage is 95%. Report valid and excluded observations separately.
5. Prompt Visibility Coverage
Prompt visibility coverage measures breadth across unique priority prompts rather than total repeated answers.
Formula:
Prompt visibility coverage = tested priority prompts with at least one qualified brand mention ÷ all tested priority prompts × 100
If the brand appears at least once for 12 of 30 tested priority prompts, visibility coverage is 40%.
This metric answers a different question from mention rate. Mention rate is sensitive to platforms and repeated runs; visibility coverage shows how much of the prompt landscape has any observed brand presence.
6. Framing Accuracy
Framing accuracy measures whether material statements about the brand match an approved fact set.
Formula:
Framing accuracy = evaluated material brand statements that match approved facts ÷ all evaluated material brand statements × 100
If reviewers evaluate 25 material statements and 21 match approved facts, framing accuracy is 84%. Classify the other four as inaccurate, contradicted, or unverifiable. Do not count unverifiable statements as accurate merely because they sound positive.
A material statement may concern category, audience, geography, capability, integration, price, certification, or another decision-relevant attribute. Use a written review guide and preserve the evidence quote.
7. Recommendation Rate
Recommendation rate measures how often the brand is actively recommended rather than merely named.
Formula:
Recommendation rate = valid answers that recommend the brand for the tested context ÷ valid recommendation-intent answers × 100
Only use prompts designed to elicit a recommendation in the denominator. A definition prompt should not count as a failed recommendation opportunity.
8. Citation-to-Mention Conversion
This diagnostic metric asks how often a mention is accompanied by a citation to the approved domain.
Formula:
Citation-to-mention conversion = answers with both a brand mention and brand-domain citation ÷ answers with a brand mention that are citation-eligible × 100
It can reveal whether the brand is known but its official pages are rarely presented as supporting sources.
Worked Metric Example
The following numbers are illustrative and do not describe an UnderAI customer.
A fictional company tests 24 prompts across three AI experiences, creating 72 planned observations. Two tests fail, leaving 70 valid answers. Fifty-eight valid answers come from citation-capable modes.
The results are:
- 21 valid answers mention the company;
- 8 citation-eligible answers cite its official domain;
- the company and four defined competitors appear 84 times in total, with 21 company appearances;
- 15 of the 24 unique prompts produce at least one company mention;
- reviewers evaluate 30 material brand statements, of which 27 match approved facts; and
- 9 of 32 valid recommendation-intent answers recommend the company.
The calculated metrics are:
| Metric | Calculation | Result |
|---|---|---|
| Mention rate | 21 ÷ 70 | 30.0% |
| Brand citation rate | 8 ÷ 58 | 13.8% |
| Share of answer | 21 ÷ 84 | 25.0% |
| Prompt test coverage | 72 ÷ 72 | 100% completed; 70 valid |
| Prompt visibility coverage | 15 ÷ 24 | 62.5% |
| Framing accuracy | 27 ÷ 30 | 90.0% |
| Recommendation rate | 9 ÷ 32 | 28.1% |
The report should also show sample sizes. A percentage without its numerator and denominator can make a small or incomplete test look more certain than it is.
Segment Metrics Before Interpreting Them
Aggregate rates can hide important differences. Segment results by:
- platform and mode;
- market and language;
- prompt intent;
- journey stage;
- audience;
- branded versus unbranded prompts;
- product or service line; and
- citation-capable versus non-citation experiences.
A visibility gain confined to branded prompts does not mean the brand gained category discovery. A high mention rate with low framing accuracy may be a risk rather than a success.
Connect Answer Metrics to Website and Business Outcomes
Answer visibility is an intermediate outcome. Where data is available, review it alongside:
- organic impressions and clicks;
- referral traffic from visible source links;
- engaged sessions and conversions;
- branded search demand;
- qualified leads and pipeline; and
- sales feedback about how buyers discovered or evaluated the brand.
Do not claim direct attribution when the platform does not expose it. Search Console reports Google's AI features within the broader Web search type rather than as a complete standalone GEO report. Analytics and CRM data can support interpretation, but they do not reveal every zero-click influence.
Compare Periods Fairly
Before calling a change an improvement, confirm that the compared periods use the same prompt versions, competitor set, platforms, modes, locales, denominator rules, and review guide. Annotate model or product changes when known.
Report both absolute counts and rates. Use repeated runs when possible, and avoid conclusions from a single answer.
Build a Decision-Oriented GEO Report
A useful report includes:
- Scope and protocol.
- Planned, completed, valid, and excluded observation counts.
- Current metrics with formulas and sample sizes.
- Segment comparisons.
- Evidence examples for gains, absences, citations, and framing errors.
- Cited-domain and competitor patterns.
- Website and business outcome context.
- Actions supported by the evidence.
- Limitations and changes to the test environment.
Use The AI Answer Authority Model to interpret the relationship between presence, evidence, context, and recommendation. Use the brand mention tracking guide to build the underlying observations.
Common Measurement Mistakes
- Mixing invalid tests into the denominator.
- Comparing citation-capable and non-citation experiences without qualification.
- Changing the competitor set between periods.
- Treating all prompts as recommendation opportunities.
- Counting positive but inaccurate framing as success.
- Reporting a proprietary composite score without the underlying metrics.
- Hiding sample sizes.
- Claiming causation from a simple before-and-after test.
- Measuring answer visibility without connecting it to business decisions.
Final Principle
GEO measurement is useful when another reviewer can understand what was tested, inspect the evidence, reproduce the calculation, and see the uncertainty. Consistent definitions matter more than a large dashboard.
