Back to blogLaptop displaying charts next to a notebook used to plan a prompt bankPractical Guides

    How to Build a Prompt Bank for Measuring Your Brand's AI Visibility

    2026-07-15·11 min read

    An AI visibility program can have an excellent dashboard and still lead to the wrong decisions. The failure often happens before any metric is calculated: in the questions selected for measurement.

    Before you run the bank, decide how you will localize the sample by country and language so every market keeps its own configuration and denominator.

    If the analysis uses generic, repetitive prompts that do not resemble how customers buy, the result does not represent the market. If too many questions include the brand name, visibility looks artificially high. If prompts change every week, trends stop being comparable.

    A prompt bank solves that problem. It is the controlled sample connecting real buyer decisions to answers from ChatGPT, Gemini, Perplexity, Claude or Google AI Overviews. This guide explains how to build one without obvious bias, how to assess its coverage and how to keep it useful over time.

    What a Prompt Bank Is and Is Not

    A prompt bank is a structured, versioned set of questions representing moments when someone might discover, compare, evaluate or choose a brand through AI.

    It is not a keyword list automatically rewritten as questions. It is not a collection of ideas produced in one brainstorming session either. Every prompt should have a business reason, an intent, a market, an audience and a rule for interpreting the answer.

    The minimum unit is not just the wording. It is this record:

    • Exact prompt.
    • Intent and decision stage.
    • Audience or buyer profile.
    • Country and language.
    • Category or product line.
    • Business value.
    • Models where it is run.
    • Creation date and version.

    With this structure, you can explain why a question belongs in the sample and prevent uncontrolled growth.

    If you productize prompt banks across an agency book, adapt them per client using the GEO guide for agencies serving multiple clients.

    Why a Keyword List Is Not Enough

    In SEO, a keyword usually represents demand that ends on a search results page. In an AI conversation, the same need can be expressed as a problem, a constraint, a comparison or a recommendation request.

    A keyword such as project management software can become very different questions:

    • "Which tools help coordinate a remote team of 30 people?"
    • "Compare three options for managing projects with external clients."
    • "Which project management software is easy to implement without a dedicated administrator?"
    • "What would you recommend for an agency that needs client-level permissions?"

    All belong to the same category, but they do not produce the same answer or carry the same value. The bank must capture those differences. For what to measure afterward, see the guide to AI visibility metrics; this article focuses on building the sample that feeds those metrics.

    The Six Blocks Your Sample Should Cover

    A balanced bank does not distribute questions randomly. It should cover six decision jobs.

    1. Category discovery. Broad questions asked before the user knows providers: "Which solutions exist for...?" They measure whether the brand enters the initial conversation.

    2. Problem and use case. The user describes a concrete need: "How can a marketing team monitor...?" The AI may recommend different categories, not just direct competitors.

    3. Comparison and shortlist. Questions asking for alternatives, advantages or a short list. They are essential for observing which brands appear together and later calculating AI Share of Voice.

    4. High-intent recommendation. These include buying context, company size, budget, industry or country: "Which tool would you recommend for...?" Their volume may be lower, but their business value is usually higher.

    5. Objections and constraints. Questions about ease of use, integrations, security, language, support or compliance. They reveal why a brand enters or leaves a recommendation.

    6. Branded diagnosis. These include your brand name: "What type of company is Brand X for?" or "What are Brand X's limitations?" They help test accuracy and framing, but must be analyzed separately. They do not measure spontaneous discovery.

    How to Build the Bank Step by Step

    1. Define the Measurement Scope

    Start by writing down the decision you want to observe. Valid measurement needs boundaries: product, market, language, audience and period. "Our brand visibility" is too broad; "accounting software recommendations for freelancers in Spain" is specific enough to support a coherent sample.

    If you have several product lines or countries, create separate banks or segments. Mixing them breaks interpretation: improvement in one category can hide decline in another.

    2. Collect Real Customer Language

    Before writing prompts, collect questions from sales calls, support tickets, site search, forms, communities, keyword research and Search Console queries. Do not copy them blindly; use them to preserve real vocabulary, constraints and use cases.

    Look especially for phrases such as "which option", "what do you recommend", "alternatives to", "for a company that", "is it worth it" or "how can I solve". They usually reveal a decision rather than simple informational curiosity.

    3. Build a Coverage Matrix

    Cross decision stages with priority segments. A simple matrix can use intent as rows and market, audience or product as columns. Each cell should contain enough prompts to avoid depending on a single wording.

    Do not force perfect symmetry. If 60 percent of revenue comes from one segment, that segment can receive more weight. Document the rule so the sample does not only reflect the preferences of whoever wrote it.

    4. Write Natural, Self-Contained Questions

    Every prompt should make sense without previous context. Avoid internal jargon, feature names customers do not use and wording that already contains the preferred answer.

    Biased prompt: "Why is Brand X the best tool for remote teams?"

    Useful prompt: "Which tools would you recommend for coordinating projects in a remote team of 30 people, and why?"

    Keep one primary intent per question. If you ask the model to compare price, security, usability, integrations and support at once, the answer becomes difficult to classify.

    5. Add Variants with a Purpose

    A variant should test a meaningful difference: country, company size, use case, constraint or buying stage. Changing which tools to what are the tools adds no coverage; it only duplicates the sample.

    Keep a common core across models for comparison. If a platform needs an adaptation, record the variant explicitly instead of silently mixing it with the base prompt.

    6. Label Intent and Value Before Measurement

    Assign each prompt an intent label and a business weight before seeing the answer. This prevents you from changing a question's importance because the result helps or hurts the brand.

    A simple scheme can use three levels: critical, important and exploratory. Critical prompts represent decisions close to purchase or reputation; exploratory prompts reveal trends. Analyze both, but do not let twenty informational questions hide five high-intent absences.

    7. Deduplicate and Run Quality Control

    Read the bank as a whole. Remove equivalent questions, balance categories and check that no competitor is disproportionately named in the wording. Another person should be able to explain what each prompt measures without asking the author.

    Before automating, run a small sample manually. If many answers interpret the wrong category or cannot be compared, fix the wording. The problem is in the design, not the dashboard.

    Minimum Template for Every Prompt

    You can manage the bank in a spreadsheet or a specialized platform. At minimum, use these fields:

    Field Example
    Stable ID EN-SaaS-Compare-012
    Prompt Which tools would you recommend for coordinating projects with external clients?
    Intent Comparison
    Audience Agency with 20-50 people
    Market and language United Kingdom / English
    Product or category Project management
    Value Critical
    Models ChatGPT, Gemini, Perplexity
    Version v1.0 - 2026-07-15
    Status Active

    The ID should not change after a minor correction. If a modification changes the intent, create a new version or prompt. This discipline makes it possible to trace what changed between measurements.

    How Many Prompts You Need

    There is no universal number. As an operating rule, a pilot can start with 30 to 50 well-distributed prompts. A brand running recurring measurement often works with 80 to 150 per relevant market or unit. Companies with several categories will need more, but they should remain segmented.

    Coverage matters more than volume. One hundred nearly identical questions do not create a more reliable sample. Before adding one, ask which matrix cell it covers and which decision it will improve.

    Start with a manageable bank, measure, identify gaps and expand through versions. The AI visibility tracker guide explains what the tool should do once this sample is properly designed.

    How to Prevent Biased Measurement

    Four controls are essential.

    Separate branded and non-branded prompts. The first diagnose how AI describes you; the second measure discovery without help. Never combine them into one percentage.

    Balance decision stages. A sample dominated by awareness can look healthy while the brand remains absent from buying comparisons.

    Do not introduce competitors arbitrarily. Naming the same two rivals repeatedly artificially increases their presence. Keep brand-specific questions in an explicit competitive block.

    Freeze a baseline. Part of the bank must remain stable across several measurements. Without that reference, you cannot distinguish real improvement from a sample change.

    Versioning and Review Frequency

    The bank should evolve without losing comparability. Use dated versions and record additions, removals, wording changes and the reason for each change.

    A useful practice is to keep 70 to 80 percent of the core stable and reserve the rest for exploration, launches or new customer questions. This is a management rule, not a statistical law; adjust it when your catalog or market changes materially.

    Review the bank monthly and before campaigns, launches or international expansion. Do not change prompts only because an answer is unfavorable. Change them when the question no longer represents a real decision or when you document a design bias.

    That versioning only holds if someone approves and records it. AI visibility measurement governance defines who may close a bank version and how a series break is documented.

    Connecting the Bank to Measurement

    The bank defines what to ask; the tool records how each model answers. From there, you can measure mention rate, position, framing, competitors, sources and stability.

    Analyze by cluster and intent first. A global average can hide the important pattern. If the brand appears in discovery but not in shortlists, the action is not to "raise the score". It is to strengthen the signals supporting a buying recommendation.

    Every prompt should be repeated under controlled conditions so differences between runs can be interpreted: see how to measure AI answer variability without biasing your benchmark before turning a single data point into a trend.

    For a ChatGPT-specific check, read how often ChatGPT recommends your brand. For multi-model tracking, keep the same prompt core and compare each channel separately.

    Mistakes That Invalidate a Prompt Bank

    1. Translating a keyword list without adding decision context.
    2. Including the brand in almost every question and calling it visibility.
    3. Adding purposeless variants to inflate sample size.
    4. Mixing countries, languages, products and audiences into one score.
    5. Changing prompts every week without preserving versions.
    6. Choosing questions after seeing which results favor the brand.
    7. Automating before manually checking that answers are comparable.

    How Mentio Applies It

    Mentio turns a question bank into recurring tracking. It runs relevant prompts across several models, detects brand and competitor mentions, records position and framing, and helps teams observe change without repeating a complete manual audit.

    Result quality still depends on design. A tool can automate execution and analysis, but the team must decide which markets, audiences and decisions matter. When both layers work, the dashboard stops being a collection of answers and becomes a decision system.

    Once the bank is versioned, use it as the base to design a GEO experiment and measure whether an action changes your visibility without confusing real change with noise.

    If you are shipping a product, derive a dated, launch-specific cohort from this bank by following the guide to monitoring a product launch in AI.

    FAQ

    What is an AI visibility prompt bank?

    It is a controlled, versioned set of questions representing how an audience discovers, compares and chooses brands through AI assistants. It provides a stable sample for measuring mentions, position, competitors and change.

    How many prompts do I need to start?

    A pilot can begin with 30 to 50 well-distributed prompts. For recurring tracking, a practical reference is 80 to 150 per relevant market or unit, prioritizing coverage over volume.

    Should I include my brand name?

    Yes, but in a separate branded block. Non-branded prompts measure spontaneous discovery; branded prompts test accuracy, positioning and objections. Combining them inflates visibility.

    Can I use the same prompts across every model?

    Keep a shared core for comparison and record adaptations required by language, location, web access or channel format. Do not silently change the wording.

    How often should I update the bank?

    Review it monthly or when products, markets or customer language change. Keep a stable baseline and manage changes through dated versions.

    How does Mentio help?

    Mentio runs the bank across several models and analyzes mentions, position, competitors, framing and change. It turns a well-designed sample into recurring, actionable measurement.

    Build the Sample First, Then Automate

    A metric is only as useful as the questions producing it. Define the market, collect real language, balance intent, remove duplicates and freeze a baseline. Then automation becomes worthwhile.

    Once the bank is in place and you are comparing platforms, the guide on how to evaluate an AI visibility tool in a 14-day trial explains how to test the bank against a real product before you sign.

    Create my first measurement in Mentio

    Want to know if AI mentions your brand?

    Discover your visibility in ChatGPT, Claude and Gemini in minutes.

    Related articles