Back to blogProfessional comparing charts to build an AI visibility industry benchmarkGEO Analytics

    AI Visibility Industry Benchmark: How to Build a Valid Comparison

    2026-08-11·17 min read

    A dashboard sorts twelve brands and places yours second. The conclusion seems immediate: "we are number two in the industry for AI visibility." Yet nobody can explain why those twelve brands entered the set, whether they serve the same buyer, whether every brand received the same questions or how many answers support each percentage.

    A valid benchmark begins before anyone queries ChatGPT, Gemini or Perplexity. It defines the population, comparable peers, observation unit, cells that must remain separate and acceptable uncertainty. Only then does it apply the same design to every cohort member.

    AI Share of Voice answers what share of observed appearances belongs to each brand within a question set. This article answers a different question: does the set of brands, prompts and executions reasonably support a comparison about an industry or peer group?

    What an Industry Benchmark Can Answer

    A sound benchmark can place a brand inside a comparable distribution during a defined window. It may show that mention rate is above the median of a specified B2B cohort, that citation rate sits inside the interquartile range, or that a gap is concentrated in evaluation, comparison or purchase intent.

    On its own, it cannot prove that the brand is the commercial leader, owns more market share, is preferred by buyers or caused a revenue change. Nor does it automatically generalize to an entire industry when the cohort contains only famous companies, platform customers or brands whose data was easy to collect.

    Write the permitted claim first, for example: "we will compare B2B SaaS brands active in Spain across four intents and three assistants during a closed window, using an identical matrix." This declares population, scope, unit and window without turning a directional result into a universal claim.

    Start With the Decision and the Estimand

    Start with the decision: detect whether a brand sits outside its peer range, whether the category has low presence, or whether a gap is concentrated in one intent.

    Then define the estimand, the quantity you intend to describe. One operational example:

    Brand mention rate in non-branded evaluation prompts, inside the Spain-Spanish-ChatGPT cell, during week 32.

    Set one primary metric and a small number of secondary measures. Switching from mention rate to average rank because the first result is inconvenient introduces post-selection. Blending presence, sentiment, position and citations into an unexplained index hides methodology choices inside one number.

    For each metric, record:

    • unit of analysis;
    • event counted as success;
    • denominator;
    • treatment of empty answers and errors;
    • time window;
    • aggregation level;
    • excluded claims.

    Define the Peer Population Before Seeing Results

    "Technology industry" is rarely a useful population. It may mix marketplaces, horizontal software, consultancies, free products and enterprise suites. Two brands compete for attention in an answer only when they solve a sufficiently similar job for a comparable audience.

    Create observable inclusion criteria. For example:

    • sells B2B software through a subscription;
    • offers a defined functional category;
    • actively markets in Spain at the cut-off date;
    • serves the same buyer and company-size range;
    • keeps its product and website active during the window;
    • can be identified unambiguously in answers.

    Add symmetric exclusions: directories that do not sell the product, discontinued brands, holdings that combine incompatible categories, managed services without comparable software or companies with no verifiable activity in the market.

    Build a candidate frame before analysis and preserve the source, date and decision for every case. When two reasonable reviewers might disagree, document the boundary rule. Do not add a brand because it "belongs" after seeing it on top, and do not remove one because it hurts your percentile.

    Turn "Same Sample" Into an Equivalent Matrix

    Giving fifty prompts to every brand does not guarantee equivalence. If twenty prompts describe one company's exclusive capability, the sample favors its catalogue. If one brand appears in branded questions and the rest only in category questions, denominators are numerically equal but conceptually different.

    The matrix should contain non-branded prompts that could reasonably return any peer under identical conditions. Stratify by intent, for example:

    Stratum Question represented Equivalence rule
    Discovery What options exist? No brand in the wording
    Evaluation Which solution fits need X? Need applies to the full cohort
    Comparison Which alternatives meet A and B? Observable, symmetric criteria
    Purchase Which provider fits context Y? Same audience, market and constraints

    The AI brand visibility prompt bank helps version the sample. For an industry benchmark, add a symmetry test: each prompt must give every eligible brand the same theoretical opportunity to be retrieved.

    Freeze Comparable Cells

    Country, language, model, retrieval mode, account, date and intent can change the outcome. Do not mix them into a mean and call it "the industry." Define a cell as the smallest combination within which a comparison remains meaningful.

    Example: Spain × Spanish × grounded Gemini × evaluation × week 32.

    Compare brands inside that cell. You may then present a panel of cells, but do not replace their bases with an opaque aggregate. The guide to measuring AI visibility by country and language explains how to separate markets; here, that separation becomes a requirement for peer comparability.

    If you need a total, fix weights before execution. Publish each cell's contribution and recalculate under reasonable alternative weights. A result that changes direction after a small model-weight adjustment is not a robust ranking.

    Balance Prompts, Runs and Exposure

    Every brand should inherit the same number of opportunities inside each stratum and cell. Generative answers vary, so one run per prompt is not enough. Apply the AI answer variability protocol within every condition and retain complete answers.

    There is no universal number of brands, prompts or repetitions. Run a pilot, estimate frequency and variability by stratum, define the smallest relevant difference, and set precision and confidence before calculating the base per cell. The NIST guidance on selecting sample sizes connects these elements. Validate the calculation with a specialist if it will move investment or reputation, and label any insufficient cell as directional. A large total cannot compensate for a weak base in the cell supporting the headline.

    Choose Metrics Without Manufacturing a Winning Index

    Start with interpretable measures:

    • Mention rate: valid answers mentioning the brand / valid answers in the cell.
    • Shortlist rate: answers including the brand in an option list / valid answers.
    • Conditional position: mean or median rank only among mentions, with coverage shown beside it.
    • Citation rate: valid answers linking a brand-owned source / answers where the model provides citations.
    • Sentiment or framing: label distribution with visible protocol and classification agreement.

    Share of Voice can accompany the benchmark, but it must not replace brand-level bases. A brand may own a high share within very few valid answers. Always disclose numerator, denominator and coverage.

    If you build a composite index, explain indicator selection, normalization, weights and aggregation. The OECD/JRC Handbook on Constructing Composite Indicators shows how those choices can alter rankings and recommends uncertainty and sensitivity analysis. Do not present convenient weights as a natural property of the industry.

    Design an Auditable Dataset

    The useful grain is one row per evaluated brand, answer and condition. Preserve benchmark_version, peer identity and inclusion rule, cell_id, prompt and classifier versions, run, original answer, validity, mention, rank and citation. An answer with three brands creates three evaluations against the same evidence. Retain the dictionary, calculation, exclusions and changes. Eurostat's reference metadata standards reinforce that concepts, methods and quality must accompany statistics for correct interpretation.

    Describe the Distribution Before Ranking Brands

    A league table from 1 to 12 removes distance and uncertainty. Two brands may occupy adjacent positions with practically indistinguishable results; another may appear first because of an extreme value in a small sample.

    At minimum, publish:

    • cohort median;
    • 25th and 75th percentiles;
    • interquartile range;
    • cohort size and observations per brand;
    • percentage of missing or invalid data;
    • uncertainty interval or band per brand;
    • distribution by cell and intent.

    NIST's box plot guidance summarizes medians, quartiles, spread and extremes, and supports group comparisons without reducing them to a mean. Use width proportional to the base where helpful and provide an accessible table beside the chart.

    You may express relative position as a cohort percentile, but write "percentile within this cohort," not "industry percentile," unless the frame covers the population. Avoid winner labels when intervals overlap or the ordering is unstable.

    Report Uncertainty and Avoid False Precision

    An observed rate is an estimate. Its interval should reflect that one hundred observations provide less precision than one thousand and that rare events produce asymmetric limits. For small proportions or limited bases, do not default to a normal approximation; use an appropriate, documented method.

    The NIST overview of confidence intervals explains that an interval describes the long-run behavior of a repeated sampling procedure, not an absolute guarantee about one result. Show the method, level, base and assumptions in the report.

    Do not declare that A beats B merely because 24.3% > 23.8%. Require the difference to exceed the relevant threshold, persist in key cells, exceed noise, survive alternative rules and have enough observations.

    Comparing many brands and metrics increases the chance of finding accidental differences. Prioritize one primary hypothesis or measure and treat the rest as exploratory, or apply a multiple-comparison procedure with statistical support.

    Separate Absence, Failure and Censoring

    "Not mentioned" is not the same as "answer unavailable." Separate a valid answer without the brand, error, refusal, truncation, ambiguous entity, censored list and unverifiable citation. Define which states enter each denominator and publish coverage by brand, cell and reason. Never impute a mention without a prior rule or turn an error into zero.

    Put the Ranking Through Sensitivity Analysis

    Before publication, recalculate under broad and strict cohorts, equal and business weights, median and trimmed mean, inclusion and exclusion of ambiguity, each model separately, and successive removal of repetitions or extreme brands.

    Record how much the metric and order move. If a brand falls from second to seventh after removing one peer or adjusting a plausible weight, the finding is sensitive. Report a position range or band instead of a rigid rank.

    The goal is not to find a configuration that confirms expectations. It is to discover which conclusions survive defensible choices.

    Example: A Twelve-Brand B2B SaaS Benchmark

    Suppose the population is B2B software in one category, active in Spain and sold to marketing teams at companies with 50 to 1,000 employees. Two consultancies and one marketplace are excluded from fifteen candidates before measurement, leaving twelve peers. The cell uses Spain, Spanish, three assistants, one week, 48 prompts balanced across four intents and five repetitions. Each brand receives 720 opportunities over the same answers: 8,640 evaluations in total.

    An illustrative result shows 24% mention rate against an 18% median, 12%/27% 25th/75th percentiles and 697 valid answers out of 720. The correct claim is that the brand sits above the median and in the upper cohort band with 96.8% coverage, with an evaluation advantage that does not persist in purchase, not that it is "second in the industry." If removing two enterprise peers or using equal weights moves it to the center, report composition dependence.

    Publish a Method Card With the Result

    The card must disclose purpose and permitted claim; population and date; candidates, inclusions and exclusions; unit; prompts, strata, models, repetitions and window; formulas, denominators and weights; errors, ambiguity and QA; summary, intervals and sensitivity; limitations, version, owner and next review. Version card and dataset together. If the cohort changes, publish a bridge or start another series. To turn the result into a citable public asset, also apply the proprietary-data methodology; this operational benchmark does not replace it.

    Checklist Before Calling the Benchmark "Industry-Level"

    • [ ] The target population was written before data collection.
    • [ ] Inclusion and exclusion rules are observable, symmetric and dated.
    • [ ] The cohort does not change after seeing the ranking and every brand shares applicable non-branded prompts.
    • [ ] Country, language, model, mode and intent remain separate cells.
    • [ ] Each cell has balanced bases or an explicit limitation.
    • [ ] Metric, denominator and error are predefined; failures and missingness are not confused with zero.
    • [ ] Median, quartiles, base, coverage and uncertainty are published.
    • [ ] Weights and any composite index are justified.
    • [ ] The result passes sensitivity checks and does not confuse cohort percentile with market share.
    • [ ] Method card and version make the calculation reproducible.

    FAQ

    What is an AI visibility industry benchmark?

    It is a comparison of brand-presence metrics in AI answers across a peer population defined before measurement. Every brand receives the same matrix of prompts, models, markets, languages, windows and repetitions, while the report discloses sample size, distribution, uncertainty and limitations.

    How should I choose brands for the cohort?

    Define the target population and observable inclusion and exclusion criteria first: category, customer, geography, business model, product scope and activity during the window. Preserve the candidate frame and cut-off date. Do not add or remove brands after seeing their results.

    How many brands and prompts do I need?

    There is no universal minimum. It depends on the population, desired precision, variability and decisions the benchmark will support. Run a pilot, set tolerable error, calculate the requirement by stratum and always disclose the base. A small sample can be directional, but not representative.

    Should I use the mean or the median?

    Start with the median, 25th and 75th percentiles, interquartile range, sample size and missing data. The mean can be added when it is useful, but it is sensitive to extreme values and skewed distributions. Do not hide the distribution behind one number.

    How should I compare countries, languages and models?

    Treat each country, language, model, mode and intent combination as a separate cell. Compare brands only within equivalent cells and report cell-level results. If you build an aggregate, set weights before measurement and show how the result changes under other reasonable weights.

    Does a benchmark prove that a brand is commercially better positioned?

    No. It describes visibility within a specific sample and window. It does not prove market share, buyer preference, causality or sales. Decisions should combine it with business metrics, customer research and a separate design for evaluating impact.

    Compare Peers, Not Incomparable Tables

    Start with one cell and one decision. Write the population, freeze the cohort, run a symmetric matrix and publish the distribution with its limits. If you cannot explain why every brand and observation is present, you do not have a benchmark yet.

    Build your AI visibility benchmark with Mentio ->

    Want to know if AI mentions your brand?

    Discover your visibility in ChatGPT, Claude and Gemini in minutes.

    Related articles