AI Search Tool Rank
All posts
By AI Search Tool Rank Teamcompetitive analysisshare of voicemeasurement

Benchmarking Competitors in AI Search: Heatmaps and Share of Voice

How to benchmark AI visibility with a fixed competitor set, model-level heatmaps, and share-of-voice denominators that remain comparable over time.

Competitor benchmarking in AI search starts with a deceptively awkward question: compared with whom? AI answers can introduce brands a company did not put in its market map. At the same time, a share-of-voice calculation changes whenever the comparison set changes. A useful benchmark must allow discovery without rewriting its denominator every week.

Heatmaps and share of voice solve different parts of that problem. Share of voice summarizes relative presence. A heatmap shows where the relative result comes from by separating brands and models. Read together, they can expose a gap that a blended visibility number would hide.

Share of voice depends on the comparison set

Promptwatch's share-of-voice documentation, reviewed August 30, 2026, defines competitive share of voice as your brand mentions divided by total mentions across the brands being compared, multiplied by 100. A brand counts once per response even if the answer repeats its name.

That makes the metric relative by design. Your mention count can rise while your share falls if competitors gain faster. It also means adding a frequently mentioned competitor expands the denominator and can lower your historical-looking share without changing any old response.

Before reporting the metric, publish the competitor set and its version. Keep a core set for period comparisons. Put newly discovered brands into a candidate list, then add them at a scheduled benchmark reset. If an emerging company must be added immediately, calculate an overlap period using both the old and new sets so readers can see the denominator effect.

Do not use every detected brand by default for an executive trend. A broad detected-brand view is useful for discovery, but it may include adjacent products that do not compete for the same buyer decision. Conversely, a handpicked list can omit the brands AI engines recommend. The two views answer different questions and should have different labels.

A heatmap exposes model-specific gaps

The competitor heatmap documentation, also reviewed August 30, 2026, describes a matrix with one row per brand and one column per AI model. Each cell contains that brand's average visibility score on that model for the selected period.

Under the documented calculation, absences count as zero in the response-level visibility average. Daily averages are then averaged across the selected period. This design combines coverage and prominence. A low cell may mean the brand seldom appears, appears weakly, or both. Open the underlying responses before deciding which explanation applies.

Two display details prevent common reading errors. First, heatmap color is relative to the current filtered view, not fixed across all dashboards. The darkest cell in one month cannot be compared by color alone with the darkest cell in another month. Compare the displayed values under the same formula. Second, a dash means the brand did not appear in any analyzed response for that model and period. A small numeric value means it appeared with low average visibility. Those states should not be collapsed.

The heatmap is most useful when read in both directions. Reading across your row reveals engine imbalance. Reading down a weak engine column shows whether the whole category is absent or competitors are visible where you are not. The latter is a narrower investigation target.

Build the benchmark before looking at winners

Start with a prompt cohort tied to buyer tasks. Category discovery, comparisons, problem diagnosis, and branded evaluation behave differently, so tag them and keep the cuts available. A benchmark dominated by branded prompts will flatter established brands and say little about unbranded discovery.

Fix the engines, locations, languages, run cadence, and observation window. Record failed runs and eligibility rules. If one engine has half the valid responses of another, an unweighted blended average may give the sparse engine more influence than intended. Report model columns first and define any later blend.

Then choose measures that retain distinct meanings:

  • mention coverage across all valid responses
  • competitive share of voice within the fixed brand set
  • visibility by model, with the scoring method disclosed
  • citation coverage for owned and third-party pages
  • position distribution rather than a single average rank

Use heatmaps for diagnosis, not causation

Suppose your Perplexity cell is low and a competitor's is high. That result establishes a measured gap for the selected prompts and period. It does not prove the competitor's latest article caused the difference.

Filter to Perplexity and inspect the answers. Which prompts produce the gap? Which domains and page types get cited? Is the competitor present because its own page is used, because a third-party comparison mentions it, or because the model names it without a citation? These routes suggest different work.

Check time as well. A one-period cell may reflect normal answer variation. The AI answer volatility guidance, updated August 1, 2026, recommends at least around three runs per prompt per platform within a rolling seven-day window and more before making single-prompt claims. A benchmark should aggregate enough repeated observations for the decision it supports.

When a content or outreach change is made, annotate it. Compare the affected prompt cluster with an unchanged cluster when possible. A later increase is consistent with the intervention, but timing alone does not establish causation because model behavior, retrieval, and competitor activity also change.

Report movement in counts as well as shares

Percentages can move while the underlying volume shrinks. If your brand records 20 mentions out of 100 competitor-set mentions, share is 20%. In a later window, 10 out of 40 produces 25%. Share rose while your mentions fell.

That arithmetic example is illustrative, not a reported market result. It shows why every share should sit beside its numerator and denominator. Include total valid responses too, since the pool of brand mentions can change independently of collection volume.

Position deserves similar treatment. An average can hide a brand that alternates between first place and absence. A distribution of appearances by position, plus the absence rate, makes that instability visible.

Turn the benchmark into a work queue

A good competitive review ends with response-level questions. Find prompts where a competitor appears and you do not. Group them by topic. Inspect cited sources, then decide whether the gap concerns owned content, offsite coverage, technical retrieval, or an inaccurate entity description.

For practical measurement, Promptwatch provides share-of-voice and competitor heatmap views alongside the responses behind them. Our Promptwatch review covers its wider feature set. The useful combination here is the ability to move from a colored cell to the prompts and citations that produced it.

Keep the benchmark reproducible. State the date window, cohort version, models, markets, competitor list, formulas, and valid response count. Then a change in share of voice can be discussed as evidence rather than as a percentage detached from how it was made.