Content Strategy
Share of AI Voice by Query Cluster: The GEO Dashboard Most Teams Actually Need
How to measure AI visibility across query clusters, not random prompts—and turn GEO from anecdote into an operating metric.
Most GEO dashboards fail because they track the wrong thing. They screenshot a few ChatGPT responses, celebrate when the brand shows up, panic when it doesn't, and call this "measurement." It isn't. The moment you stop measuring AI visibility prompt-by-prompt and start measuring it by query cluster, GEO becomes a strategy problem instead of a screenshot habit.
The right question is not "Did we show up for this one prompt?" It is: "How visible are we across the clusters of questions that matter commercially?"
This article explains why query clusters are the missing unit of GEO measurement, how to define Share of AI Voice (SAIV) properly, how to separate mentions from citations from recommendations, and how to build a GEO dashboard that actually supports decisions.
TL;DR
- Share of AI Voice (SAIV) measures how often your brand appears—by mention, citation, or recommendation—in AI-generated answers for a defined query set.
- Single-prompt tracking is unreliable. AI answers are nondeterministic. The same prompt produces different results minutes apart.
- The query cluster is the right unit. A cluster of 8–15 related prompts smooths random variation and maps to real user intent.
- Mentions ≠ citations ≠ recommendations. These are different levels of visibility with different commercial value. Never aggregate them into one number.
- A useful GEO dashboard has 6 views: cluster-level SAIV summary, citation source map, competitor dominance, prompt volatility, intent-weighted visibility, and outcome-proxy layer.
- GEO is not separate from SEO. Query clusters map to topic clusters. Cited URLs reveal which content formats win. The GEO dashboard should drive your editorial roadmap.
- Operationalize monthly. Assign cluster owners, run monthly GEO reviews, and connect SAIV insights to content prioritization.
Definitions: The Language of GEO Measurement
Before building anything, agree on terms. Ambiguous definitions produce useless dashboards.
| Term | Definition |
|---|---|
| Share of AI Voice (SAIV) | The proportion of AI-generated answers—for a defined query set—that mention, cite, or recommend your brand. |
| Query Cluster | A grouped set of 8–15 prompts expressing similar user intent. The fundamental unit of GEO measurement. |
| Mention | Your brand name appears in the AI answer. Awareness signal only. |
| Citation | Your domain or specific page is linked as a source. Authority and trust signal. |
| Recommendation | Your brand is listed as a preferred option among alternatives. Commercial preference signal. |
| GEO | Generative Engine Optimization. Optimizing content for visibility within AI-generated answers. |
| AEO | Answer Engine Optimization. Often used interchangeably with GEO; focuses on direct-answer interfaces. |
| Prompt Volatility | The degree to which AI answers vary for the same prompt across runs. High volatility = unreliable single-prompt measurement. |
| Cluster-Level Visibility | Your aggregated SAIV score across all prompts within a cluster. Smooths prompt-level noise. |
| Citation Share | The percentage of cited sources in a cluster that point to your domain vs. competitors. |
| Recommendation Share | The percentage of AI recommendations in a cluster where your brand is included as a preferred option. |
Why Most GEO Measurement Is Wrong
Most teams attempting GEO measurement make the same mistakes. The result is noise dressed up as insight.
Common Failure Modes
- Tracking a handful of random prompts. You pick 5–10 questions that feel important, check them once, and extrapolate. This tells you almost nothing. AI answers are nondeterministic—the same prompt returns different results minutes apart.
- Assuming prompt results are stable. They are not. Run the same query 10 times and you'll get varying citations, different brand mentions, and inconsistent recommendations. Single-prompt snapshots are noise.
- Mixing low-intent and high-intent prompts. "What is marketing attribution?" and "best marketing attribution tool for SaaS" have completely different commercial value. Averaging them together destroys insight.
- Treating all appearances equally. A casual mention ("some companies like Acme...") is not the same as a citation ("according to Acme's research...") or a recommendation ("top options include Acme"). These must be scored differently.
- Counting brand mentions as commercial wins. You can be mentioned 100 times and never be recommended. Mentions are awareness. Recommendations are preference. Conflating them creates false confidence.
- Ignoring competitor share. Your absolute SAIV number means nothing without knowing who else appears—and how often. A 40% mention share sounds good until you realize one competitor has 60%.
- No business weighting. Not all clusters matter equally. A cluster about pricing comparisons might be 10x more commercially valuable than a cluster about industry definitions. Unweighted dashboards create false prioritization.
These approaches create dashboards that feel productive but support zero decisions. They are measurement theater.
Why the Query Cluster Is the Right Unit of Analysis
A single prompt is the wrong unit of GEO measurement for the same reason a single keyword is the wrong unit of SEO strategy. It's too granular, too noisy, and too disconnected from how people actually research.
1. A Single Prompt Is Unstable
AI answers are probabilistic, not deterministic. The same prompt produces varying output across runs, models, and contexts. You cannot build a measurement system on a unit that changes every time you measure it.
2. A Cluster Smooths Random Variation
When you aggregate 8–15 related prompts, random answer volatility cancels out. Your cluster-level SAIV score becomes stable enough to track trends over time and compare periods meaningfully.
3. Clusters Reflect Real User Intent
Users don't ask one question. They explore a topic through multiple related queries. A cluster mirrors this research behavior and tells you whether you're visible across an intent space—not just for one phrasing.
4. Clusters Enable Pattern Detection
At the cluster level, you can spot structural patterns: "We're always mentioned but never cited in the attribution cluster" or "Competitor X dominates the CRM selection cluster." Prompt-level data cannot reveal these patterns.
5. Clusters Align with Content Strategy
Query clusters map directly to topic clusters and content pillars. This means your GEO dashboard can directly inform which content themes need investment—making it a planning tool, not just a scorecard.
What a Query Cluster Looks Like in Practice
A cluster is a grouped set of prompts expressing similar intent. Here are examples:
- Attribution & Measurement: "How to measure marketing attribution in 2026," "best attribution models for B2B," "incrementality testing vs attribution," "marketing attribution without cookies"
- CRM Selection: "best CRM for SaaS startups," "HubSpot vs Salesforce for mid-market," "CRM evaluation criteria," "CRM implementation checklist"
- GEO/AEO Measurement: "how to measure GEO visibility," "share of AI voice definition," "AI search visibility metrics," "GEO vs SEO measurement"
- Lifecycle Orchestration: "lifecycle marketing best practices," "email frequency capping strategy," "cross-channel orchestration architecture," "lifecycle control tower"
- Vendor Evaluation: "best marketing analytics tools," "CDP comparison 2026," "top conversion optimization tools," "marketing attribution platforms ranked"
With cluster-level measurement, teams can answer: Where are we visible? Where do competitors dominate? Which content themes are underperforming? Where are the commercial visibility gaps?
The 5-Layer Share of AI Voice Model
A rigorous GEO measurement framework needs five dimensions. Miss any one, and your dashboard produces incomplete or misleading signals.
Layer 1: Query Cluster
The thematic grouping. Each cluster contains 8–15 prompts expressing similar intent across different phrasings, angles, and specificity levels.
What teams miss: Clusters that are too broad (mixing informational and evaluative prompts) or too narrow (splitting what should be one cluster into three).
Layer 2: Prompt Set & Phrasing Coverage
The individual prompts within each cluster. Include informational ("what is X"), comparative ("X vs Y"), evaluative ("best X for Y"), and action-oriented ("how to implement X") phrasings.
What teams miss: Over-indexing on informational queries and under-indexing on comparative/evaluative queries where commercial value is highest.
Layer 3: Appearance Type (Mention / Citation / Recommendation)
For each prompt result, classify the appearance: Was the brand mentioned casually? Was the domain cited as a source? Was the brand recommended as a preferred option?
What teams miss: Lumping all three into "we showed up." A mention is not a citation. A citation is not a recommendation.
Layer 4: Competitor Capture
For each cluster, map which competitors appear, how often, and in what capacity. Your SAIV is only meaningful relative to the competitive landscape.
What teams miss: Tracking only their own visibility without knowing who else dominates each cluster.
Layer 5: Intent Weight & Business Value
Not all clusters matter equally. Weight clusters by commercial importance: a pricing comparison cluster might be 3x more valuable than a glossary cluster. Apply weights to produce a business-aligned SAIV score.
What teams miss: Treating all clusters equally, which makes the dashboard decorative rather than prioritizing.
The Metrics That Matter in a GEO Dashboard
Not every metric is equally useful. Here are the metrics that actually support decisions, organized from awareness to business value:
| Metric | What It Means | Why It Matters | Common Misuse |
|---|---|---|---|
| Mention Share | % of cluster prompts where brand is named | Awareness baseline | Treated as a win when it's just awareness |
| Citation Share | % of cited sources pointing to your domain | Authority/trust signal | Not tracked separately from mentions |
| Recommendation Share | % of recommendations including your brand | Closest to purchase intent | Confused with citations |
| Competitor Share | Competitor visibility per cluster | Relative positioning | Tracked only for self, ignoring competitors |
| Unique Cited URLs | Which pages are being cited | Content format/quality signal | Not mapped back to content strategy |
| Cluster Volatility | How much results vary across runs | Measurement confidence | Ignored—leads to overreaction to noise |
| Priority Cluster Coverage | % of high-value clusters with visibility | Strategic gap analysis | All clusters treated equally |
| Weighted SAIV | Business-weighted composite score | Executive summary metric | Used without understanding the inputs |
Leading indicators: Mention share, citation share, cluster volatility. These move before business outcomes shift.
Closer to business value: Recommendation share, weighted SAIV, priority cluster coverage. These correlate with branded search, inbound quality, and pipeline influence.
How to Build the GEO Dashboard
A practical GEO dashboard needs six views. Each one answers a different strategic question.
View 1: Cluster-Level SAIV Summary
What it shows: SAIV score for each cluster, broken into mention/citation/recommendation. Trended over time.
Decision it supports: "Which clusters need content investment?" and "Where are we gaining or losing ground?"
View 2: Citation Source Map
What it shows: Which specific URLs from your site are being cited, in which clusters, how often.
Decision it supports: "Which content is working?" and "Which pages should we update or promote?"
View 3: Competitor Dominance View
What it shows: Side-by-side visibility of top competitors per cluster. Who appears most? In what capacity?
Decision it supports: "Where are competitors winning?" and "Which clusters should we contest?"
View 4: Prompt Volatility View
What it shows: How much variation exists across runs for each cluster. High volatility = low confidence in trends.
Decision it supports: "Can I trust this cluster's data?" and "Where should I increase prompt coverage?"
View 5: Intent-Weighted Visibility
What it shows: SAIV weighted by commercial importance. High-value evaluative clusters weighted higher than informational ones.
Decision it supports: "Where should we prioritize content creation for maximum business impact?"
View 6: Outcome-Proxy Layer
What it shows: Correlation between SAIV trends and proxy outcomes—branded search volume, direct traffic, assisted conversions, inbound lead quality.
Decision it supports: "Is GEO visibility actually helping the business?"
How to Design Query Clusters
The quality of your GEO dashboard depends entirely on the quality of your clusters. Here is how to build them properly.
Choosing Clusters
Start with 5–7 clusters mapped to your most important content themes and commercial priorities. Each cluster should represent a distinct intent space your business needs to be visible in.
Prompts Per Cluster
Use 8–15 prompts per cluster. Fewer than 8 and you don't have enough data to smooth volatility. More than 15 and you risk diluting the cluster's focus.
Prompt Diversity Rules
Each cluster should include prompts across four intent types:
- Informational: "What is marketing attribution?"
- Comparative: "Marketing attribution vs incrementality testing"
- Evaluative: "Best marketing attribution tools for B2B SaaS"
- Action-oriented: "How to implement incrementality testing"
Weighting Clusters
Assign each cluster a business weight from 1 (low) to 5 (high) based on:
- Revenue relevance of the topic
- Search volume of related queries
- Strategic importance to positioning
- Competitive intensity in the space
Sample Cluster Map
| Cluster | Prompts | Business Weight | Primary Intent |
|---|---|---|---|
| Attribution & Measurement | 12 | 5 (critical) | Evaluative + Action |
| GEO / AEO Visibility | 10 | 4 (high) | Informational + Evaluative |
| CDP vs CRM vs Warehouse | 12 | 5 (critical) | Comparative |
| Lifecycle Orchestration | 10 | 3 (medium) | Action + Evaluative |
| Conversion & Funnel Design | 8 | 4 (high) | Action |
| Vendor Selection / Best Tools | 15 | 5 (critical) | Evaluative |
Why Mentions, Citations, and Recommendations Must Be Separated
This is the most common source of false confidence in GEO measurement. Teams celebrate "showing up" without understanding how they showed up.
Mention = Awareness Signal
The AI names your brand in passing. "Some companies in this space include Acme, BrandX, and Others." This means the model has seen your brand in training data. It does not mean the model trusts you or prefers you.
Citation = Authority Signal
The AI links to your domain or references your content as a source. "According to Acme's research on attribution modeling..." This means the model treats your content as a credible reference. It's an authority signal, not a preference signal.
Recommendation = Commercial Preference Signal
The AI lists your brand as a recommended option. "For mid-market SaaS companies, strong options include Acme (best for data-driven teams)..." This is the highest-value appearance type. It directly influences purchase consideration.
Suggested Scoring
- Mention: 1 point
- Citation: 3 points
- Recommendation: 5 points
This weighting ensures your composite SAIV score reflects commercial value, not just brand awareness.
How GEO Connects to SEO
GEO is not separate from SEO. It is an extension of it—and a feedback loop.
- Strong rankings often support citations. Pages that rank well in Google are more likely to appear in AI training data and retrieval-augmented generation (RAG) systems. But the correlation is imperfect—a lower-ranking page with better structure may be cited more.
- Structured, specific content performs better in both. The same content qualities that earn Featured Snippets (clear definitions, structured lists, specific data) also earn AI citations.
- Query clusters map to topic clusters. If you already run topic-cluster SEO, your GEO clusters can mirror those themes. One measurement system, two channels of visibility.
- Cited URLs reveal content format winners. When you see which pages are cited, you learn what content formats the AI prefers. Comparison pages? Definition articles? Framework posts? This should drive your editorial roadmap.
The GEO dashboard should inform your content strategy: which topics to update, what comparison pages to create, which glossary entries to write, and which implementation guides to publish.
Practical Examples
Example 1: Random Prompt Tracking (What Fails)
A B2B SaaS company picks 20 prompts across different topics, checks them once a month, and records "appeared" or "did not appear." Results: 8 out of 20 appearances. They report "40% SAIV" to leadership.
Why this fails: The 20 prompts span six unrelated topics. Some are informational, some commercial. They mix mentions with citations. They didn't check competitors. They didn't run prompts multiple times. Next month, different prompts hit, and "SAIV" swings to 55%. Leadership thinks GEO is working when nothing changed.
Example 2: Cluster-Based Measurement (What Works)
The same company reorganizes into 5 clusters with 10 prompts each. They run each prompt 3 times. They separate mentions, citations, and recommendations. They track competitors.
Findings:
- Attribution cluster: 60% mention share, but only 15% citation share and 5% recommendation share. They're known but not trusted.
- CRM selection cluster: 20% mention share, but 40% citation share. Their comparison article is performing well as a source.
- Vendor evaluation cluster: Competitor X dominates with 70% recommendation share. Urgent gap.
Now the team knows exactly where to invest: build authority content for the attribution cluster, protect the CRM citation, and create a competing comparison page for vendor evaluation.
Example 3: Dashboard-Driven Content Prioritization
A content team reviews the monthly GEO dashboard and sees that the "CDP vs CRM vs Warehouse" cluster has low citation share (10%) despite being weighted as a critical cluster.
Action: They build a stronger comparison article with specific evaluation criteria, pricing benchmarks, and implementation timelines. They add FAQ schema and structured definitions. Three months later, citation share increases to 35%. The comparison page becomes the team's highest-performing content asset.
Operational Playbook: Running GEO as a Real Operating System
A dashboard without a cadence is a decoration. Here is how to operationalize cluster-based GEO measurement.
Monthly GEO Review
30–45 minutes. Review cluster-level SAIV trends, competitor shifts, and citation source changes. Output: 2–3 content actions for the next sprint.
Cluster Ownership
Assign each cluster to a content owner or topic lead. They are responsible for prompt maintenance, trend interpretation, and content recommendations for their cluster.
Prioritization Framework
Focus on clusters with: (1) high business weight + low SAIV, (2) declining citation share, (3) competitor dominance that threatens positioning. Ignore high-SAIV, low-weight clusters.
Link to Editorial Roadmap
Every monthly review should output specific content recommendations: update this article, create this comparison page, add FAQ schema to this post, publish this new framework. GEO insights should feed directly into your content calendar.
Connect to SEO Reporting
Run GEO and SEO reviews together. Correlate SAIV changes with organic traffic, branded search volume, and ranking shifts. Over time, build an understanding of which levers move both channels.
Avoid Dashboard Bloat
Start with 5–7 clusters and 6 dashboard views. Add complexity only when you've outgrown the current setup. The best GEO dashboard is the one your team actually reviews monthly.
Common Mistakes to Avoid
- Measuring prompts, not clusters. You get noise instead of signal.
- Counting mentions as wins. Mentions are awareness. Only recommendations indicate preference.
- No business weighting. Treating all clusters equally makes the dashboard decorative.
- No competitor view. Your SAIV means nothing without relative positioning.
- No source page mapping. If you don't know which pages are cited, you can't improve them.
- No content action loop. A dashboard that doesn't produce editorial actions is a vanity project.
- Overfitting to unstable results. Prompt-level swings are noise. Only cluster-level trends matter.
- Building a vanity dashboard. If no one makes a different decision because of your dashboard, it's not a dashboard—it's a screensaver.
Minimum Viable GEO Dashboard Checklist
- ☐ 5–7 query clusters defined and documented
- ☐ 8–15 prompts per cluster with intent diversity
- ☐ Business weights assigned to each cluster (1–5)
- ☐ Mentions, citations, and recommendations tracked separately
- ☐ Competitor visibility captured per cluster
- ☐ Prompts run 3+ times per measurement period to reduce volatility
- ☐ Cluster-level SAIV summary view built
- ☐ Citation source map showing which URLs are cited
- ☐ Competitor dominance view per cluster
- ☐ Monthly review cadence scheduled
- ☐ Cluster owners assigned
- ☐ Content action loop connected to editorial roadmap
- ☐ Outcome proxies (branded search, direct traffic) tracked alongside SAIV
Frequently Asked Questions
What is Share of AI Voice?
Share of AI Voice (SAIV) is the proportion of AI-generated answers—for a defined query set—that mention your brand, cite your domain, or recommend you as a preferred option. It measures your visibility in AI answer engines like ChatGPT, Perplexity, and Google AI Overviews.
How is Share of AI Voice different from SEO visibility?
SEO visibility measures ranking positions and click-through rates in traditional search results. SAIV measures inclusion in AI-generated answers—whether your brand is mentioned, cited, or recommended. You can have zero SEO visibility and high SAIV, or vice versa.
Why should I measure by query cluster instead of individual prompts?
Individual prompts produce nondeterministic results—the same question returns different answers across runs. Clusters of 8–15 related prompts smooth this volatility and reveal stable, actionable patterns about your visibility across intent spaces.
What is the difference between a mention and a citation?
A mention means the AI names your brand in passing (awareness). A citation means the AI links to your domain or references your content as a source (authority). A citation carries roughly 3x the strategic value of a mention.
What is recommendation share?
Recommendation share is the percentage of AI responses in a cluster where your brand is listed as a preferred option among alternatives. It's the highest-value appearance type and the closest proxy to purchase influence.
How many prompts should be in a query cluster?
8–15 prompts per cluster. Fewer than 8 doesn't provide enough data to smooth prompt-level volatility. More than 15 risks diluting the cluster's intent focus. Include informational, comparative, evaluative, and action-oriented phrasings.
How often should I measure GEO visibility?
Monthly is the recommended cadence. AI models update frequently enough that monthly captures real trends, but not so frequently that you overreact to noise. Run each prompt 3+ times per measurement period.
Can small sites compete on Share of AI Voice?
Yes. AI retrieval systems evaluate passage quality, not just domain authority. A niche site with clear, specific, well-structured content can earn citations and recommendations over larger, less structured competitors.
How do I choose the right query clusters?
Start with your most commercially important topics—the themes where AI visibility directly influences buyer decisions. Map clusters to your existing content pillars or topic clusters. Weight by revenue relevance and competitive intensity.
How do I know if GEO visibility is helping the business?
Track outcome proxies alongside SAIV: branded search volume, direct traffic trends, inbound lead quality, and assisted conversions. Over time, correlate SAIV improvements with these business signals to build confidence in the channel.
What tools do I need to build a GEO dashboard?
At minimum: a systematic way to run prompts across AI engines, a spreadsheet or database to classify results (mention/citation/recommendation), and a reporting layer. Dedicated tools like Otterly, Profound, or custom scripts can automate prompt execution.
Should I track multiple AI engines or just one?
Track at least 2–3 engines (e.g., ChatGPT, Perplexity, Google AI Overviews). Each engine uses different retrieval logic and training data. A single-engine view gives you an incomplete picture of AI visibility.
Conclusion
GEO becomes useful only when you measure at the right level. The individual prompt is too noisy. The cluster is the right unit.
The best GEO dashboard is a prioritization system, not a vanity scorecard. It tells you where you're visible, where you're not, where competitors dominate, and where content investment will move the needle. It separates awareness (mentions) from authority (citations) from preference (recommendations). It connects to business outcomes through weighted scoring and outcome proxies.
Most teams will continue to screenshot ChatGPT responses and call it measurement. The teams that build cluster-based dashboards will have something their competitors don't: a systematic, defensible, decision-grade understanding of their AI visibility.
If you measure prompts, you get noise. If you measure clusters, you get strategy.