Back to blogInternational team comparing AI visibility data by country and languageGEO Analytics

    How to Measure AI Visibility by Country and Language Without Mixing Markets

    2026-07-27·13 min read

    A brand can appear in a Spanish answer from Spain and disappear for an equivalent question from Mexico. It can also be visible in English for the United States but absent in the United Kingdom. If all those observations end up in one rate, the number looks precise but describes no real market.

    The fix is not adding a column called "country" at the end of the analysis. Measurement must be designed so every country-language combination retains its own sample, configuration, denominator and baseline.

    This guide begins after you have built an AI brand visibility prompt bank. We will take defined intents and create a localized system that answers: are we improving in each market under comparable conditions?

    The error: turning different markets into one average

    Imagine 100 runs:

    • Spain in Spanish: 40 runs, 20 mentions.
    • Mexico in Spanish: 20 runs, 4 mentions.
    • United States in English: 40 runs, 8 mentions.

    The aggregate rate is 32 / 100 = 32%. That number hides three realities:

    • Spain: 50%.
    • Mexico: 20%.
    • United States: 20%.

    If Spain grows next month while Mexico declines, the average may improve even though a priority market gets worse. The reverse can also happen: a small market with an unstable sample moves the total too much.

    A global average is interpretable only after every market has been measured separately and its weights are known. It must never replace local views.

    Define the unit: the country x language cell

    The minimum unit is a measurement cell. At minimum it contains:

    target country x query language x model x mode

    Examples:

    • ES-es | ChatGPT | web search enabled
    • MX-es | Gemini | clean account
    • US-en | Perplexity | web
    • US-es | ChatGPT | web search enabled

    Country and language are not equivalent. US-es may represent Spanish-speaking buyers in the United States, while MX-es represents another market using the same language. Likewise, CA-en and CA-fr share a country but not a language audience.

    Every cell needs:

    1. Objective and audience.
    2. Included intent list.
    3. Local prompt variants.
    4. Execution configuration.
    5. Acceptance and exclusion rules.
    6. Its own denominator.
    7. Baseline and version.
    8. QA owner.

    Step 1: Decide which markets deserve a cell

    Do not open twenty countries because a tool displays a flag. Start with one decision question:

    In which markets do we need to know whether a person can discover, compare or choose our offer through AI during the next quarter?

    Prioritize with operating evidence:

    Signal Question
    Revenue or pipeline Does the market already generate material business?
    Investment Is there a local campaign, team or launch?
    Availability Can the offer be bought and delivered there?
    Risk Would an incorrect description have legal or commercial impact?
    Opportunity Is there a priority category or audience without a baseline?

    A cell with no decision attached creates maintenance, not information. Record excluded markets and their review date too.

    The GEO opportunity for businesses in different markets explains why regional strategy matters. This protocol solves another question: how to measure countries without treating different audiences as a single market.

    Step 2: Create a master intent set

    Market comparison requires conceptual equivalence, not identical sentences.

    Assign a stable ID to every intent:

    ID Intent Stage Expected outcome
    INT-01 Discover category solutions Discovery List of relevant options
    INT-02 Compare two approaches Evaluation Differences and criteria
    INT-03 Solve a specific problem Consideration Applicable recommendation
    INT-04 Validate trust Decision Evidence, limits and alternatives

    The ID connects local variants without claiming they are the same phrase. If an intent does not exist in a market, mark it not applicable; do not invent a translation to fill the table.

    Keep a common core for comparison and a local block for exclusive needs:

    • Comparable core: intents that exist in every selected cell.
    • Local block: questions relevant only to a market's regulation, channel, product or behavior.

    Calculate international comparisons only from the comparable core. Report the local block within its market.

    Step 3: Localize by intent, not translation

    A literal translation can change naturalness, category or the competitor set a user expects.

    For every local variant, record:

    Field Example
    Intent ID INT-01
    Cell MX-es
    Prompt ¿Qué plataformas sirven para medir visibilidad de marca en respuestas de IA en México?
    Category term plataformas de visibilidad en IA
    Local context Available to Mexican companies
    Expected entities Category, no forced brands
    Local reviewer Mexico owner
    Status Approved

    A local reviewer should check:

    • The phrase would actually be used in that market.
    • It contains no localism imported from another country.
    • The category and offer exist there.
    • Currency, regulation and availability are correct.
    • The question does not introduce the answer.
    • It retains the master intent.

    Do not force the country name into every prompt. Include it when a person would need it to obtain a local answer; otherwise, location and language belong in the configuration.

    Step 4: Freeze the configuration sheet

    Results can change because of more than text. Create one sheet per cell and version:

    text
    Cell ID:
    Target country:
    Query language:
    Expected response language:
    Model and version:
    Mode / web access:
    Account state:
    Location method:
    Session rule:
    Execution window:
    Prompt-bank version:
    Entity and competitor rules:
    Owner:

    Keep these stable during a comparison:

    • Model and mode.
    • Session and memory state.
    • Account type when it affects features.
    • Method used to set location.
    • Time window and cadence.
    • Prompt version.
    • Matching rules for brand, products and competitors.

    Do not describe a run as "Mexico" if you changed only the prompt language. The target country needs a verifiable method in the tool or environment. If you cannot confirm it, label the observation by language, not country.

    Step 5: Execute within each cell

    Randomness still exists inside a local configuration. Apply the AI answer variability protocol separately within each cell.

    The correct order is:

    1. Freeze the cell.
    2. Run approved variants.
    3. Repeat under the same configuration.
    4. Record responses and exclusions.
    5. Calculate local metrics.
    6. Compare with that cell's baseline.
    7. Only then build a multi-market view.

    Do not compensate for missing Mexico runs by adding more Spain runs. If a cell does not reach the agreed minimum coverage, label it incomplete.

    Minimum dataset for traceability

    Each row should represent one run:

    Field Content
    run_id Unique ID
    cell_id Country, language, model and mode
    intent_id Master intent
    prompt_variant_id Local text and version
    target_country Declared country
    query_language Requested language
    response_language Observed language
    model Model and mode
    executed_at Date and time
    accepted Yes / no
    exclusion_reason Reason when omitted from metrics
    brand_mentioned Yes / no
    position Position defined by the protocol
    competitors Detected entities
    sources URLs or domains when available
    raw_response Complete evidence

    Store the entity-dictionary version too. The same brand may use different aliases by country, and a product may not be available in every cell.

    Calculate denominators by cell

    The local mention rate is:

    accepted runs with a mention / accepted runs in the cell

    Do not use as the denominator:

    • All planned prompts when some were not run.
    • Answers in another language.
    • Runs excluded for incorrect configuration.
    • Intents that do not apply to the market.
    • Repetitions added only to one cell.

    Always publish numerator, denominator and coverage:

    MX-es: 18 mentions / 60 accepted runs; 95% plan coverage

    Apply the same rule to Share of Voice, position, citations and presence by intent. The general AI visibility measurement guide defines the metrics; the requirement here is that each one keeps its local universe.

    Build an aggregate only with explicit weights

    Three views are useful:

    1. By cell: the main operating view.
    2. Unweighted mean: gives every complete cell equal weight.
    3. Weighted index: reflects a documented business priority.

    Example:

    Cell Mention rate Business weight
    ES-es 50% 50%
    MX-es 20% 30%
    US-en 20% 20%

    Weighted index:

    (0.50 x 0.50) + (0.20 x 0.30) + (0.20 x 0.20) = 35%

    Weighting does not improve data quality. It only answers a business question. Version the weights, show the cell view and avoid changing them between periods without recalculating history.

    Compare deltas, not absolute positions across markets

    A higher rate in Spain than Mexico does not prove the Spain team performs better. Cells may differ in competition, sources, category maturity and model behavior.

    More defensible comparisons are:

    • ES-es current against its own baseline.
    • MX-es current against its own baseline.
    • Change in a shared intent within each cell.
    • Difference between treated and control cohorts within the same market.
    • Coverage and data quality by cell.

    Then compare deltas:

    Cell Baseline Current Delta
    ES-es 42% 50% +8 pp
    MX-es 24% 20% -4 pp
    US-en 18% 20% +2 pp

    The correct interpretation is not "Spain wins." It is "Spain improves against its starting condition, Mexico declines and the United States remains close to baseline."

    Example: Spain, Mexico and the United States

    A SaaS company operates in Spain and Mexico and begins selling in the United States. It defines:

    • ES-es: comparable core plus questions about integrations used in Spain.
    • MX-es: the same core by intent, with Mexican terminology and availability.
    • US-en: equivalent core in English, with US competitors and categories.
    • US-es: an exploratory cell for Spanish-speaking buyers, reported separately.

    During QA, it finds three MX-es answers in English and two US-en prompts containing a product unavailable there. It excludes those runs, fixes the next version and does not fill the gaps with data from another cell.

    The report shows four local results, coverage, exclusions and deltas. Leadership also gets a weighted index with approved commercial weights, but can open each market and see what explains the total.

    QA to detect cross-market contamination

    Before accepting a run, verify:

    1. The target country is confirmed by the agreed method.
    2. The prompt matches the cell variant and version.
    3. The answer uses the expected language or an explicit exception rule.
    4. Product, price, currency and availability fit the market.
    5. Compared competitors operate there.
    6. The session retains no context from another cell.
    7. Model and mode match the configuration sheet.
    8. The run falls within the defined time window.
    9. Alias and entity rules are correct.
    10. Complete evidence can be reviewed.

    Include local control prompts. For example, use an availability question that should produce different answers in Spain and Mexico. If both cells consistently return the same offer, review configuration before interpreting visibility.

    Minimum dashboard by market

    The overview should answer which market needs attention first:

    Cell Coverage Mention Share of Voice Delta Status
    ES-es 100% 50% 31% +8 pp Improving
    MX-es 95% 20% 14% -4 pp Review
    US-en 100% 20% 18% +2 pp Stable

    When a cell is opened, show:

    • Local intents and prompts.
    • Repetition distribution.
    • Brand, position, competitors and sources.
    • Excluded answers and reasons.
    • Baseline and version changes.
    • Owner and next action.

    For leadership, move the weighted index and decisions into an AI visibility executive report. Keep local detail available so the summary never turns an average into an explanation.

    Mistakes that invalidate international measurement

    1. Using language as a substitute for country.
    2. Translating prompts literally without validating local intent.
    3. Mixing answers from several markets into one denominator.
    4. Comparing raw mentions with different sample sizes.
    5. Changing business weights without recalculating history.
    6. Allowing products or competitors unavailable in the cell.
    7. Running several cells in a session with shared memory.
    8. Hiding excluded answers or incomplete coverage.
    9. Claiming a location the tool cannot confirm.
    10. Presenting the aggregate without allowing each market to be opened.

    FAQ

    Why should I not mix countries in one AI visibility metric?

    Each country and language can have different prompts, sources, competitors, availability and behavior. Mixing them lets an improvement in a large market hide a decline elsewhere and leaves the denominator describing no specific condition.

    What is a country-language cell?

    It is the minimum unit of localized measurement. It combines a target country and query language with a documented model, account, location, prompt and execution-window configuration. Its results keep their own denominator and baseline.

    Should I translate the same prompts literally?

    No. Preserve the intent and funnel stage, but adapt vocabulary, category, location and real market alternatives. A literal translation may sound unnatural or represent a different buying decision.

    How do I compare markets with different sample sizes?

    Compare rates within each cell and apply explicit business weights only for a global view. Do not add raw mentions. Keep an unweighted view too, so a commercial weight cannot hide data-quality problems.

    Can I compare Spanish in Spain with Spanish in Mexico?

    Yes, if they remain separate cells with intent-equivalent prompts, controlled configuration and their own baselines. A shared language does not make them one sample because terms, competitors, availability and local context differ.

    Which controls prevent contamination in an international benchmark?

    Record country, requested and observed language, prompt variant, model, mode, account, location, date and response. Add local control tests and exclude runs with the wrong language, unconfirmed location or different configuration.

    That market separation also constrains brand comparisons: the AI visibility industry benchmark requires comparing peers only inside equivalent country, language, model and intent cells.

    Separate first; compare second

    Multi-market measurement is useful when every market can explain its own result. Design cells, preserve local baselines and denominators, and leave aggregation until the end. A global improvement will then never hide which country advanced, which declined or where evidence is missing.

    Check your AI visibility with Mentio ->

    Want to know if AI mentions your brand?

    Discover your visibility in ChatGPT, Claude and Gemini in minutes.

    Related articles