How to Weight an AI Prompt Bank by Demand and Business Value
A prompt bank does not represent a queue of identical jobs. A discovery question may shape an entire category, a comparison may appear just before a shortlist, and a risk question may block a purchase even when it has little Google exposure. Giving every row the same weight means the average describes the size of your list, not necessarily the importance of the decisions you want to protect.
Weighting states what each observation represents, which signals justify its importance and which denominator remains stable while you compare waves. The result should be auditable: another person should be able to rebuild the weight and understand what would change if the business priority were different.
This guide starts with an existing prompt bank. It does not teach you to invent questions or choose a monitoring vendor. It shows how to build a transparent rule that orders a sample by observable demand, economic value, strategic importance and data confidence. The method and numbers are a hypothetical Mentio example, not an official Google metric or a prediction of how a model will answer.
What weighting is meant to solve
A simple average has one advantage: it is easy to explain. It can also hide three problems:
- One family with many variants can dominate because it has more rows.
- A low-frequency question can be critical to a commercial decision.
- Missing business data can accidentally become zero and push a row to the bottom.
Before assigning weights, write the decision the metric must support. “Improve visibility” is too broad. “Decide which prompt family deserves a content test this month” is operational. The rule should serve that decision rather than pretending to summarize the whole market.
Separate four concepts:
- Unit: what a row represents: a question, an intent family or a market-model cell.
- Outcome: what you observe: a mention, position, attribute, citation or condition.
- Weight: how much the unit influences the aggregate indicator.
- Eligibility: whether the row has a comparable observation in that wave and may enter the denominator.
A weight does not make an observation more true. It only changes how much it matters for the aggregate decision. Keep an unweighted rate as a control so you can see whether a family gains influence by design rather than by real coverage.
Start with the right unit
Do not assign a weight before deduplicating. If five phrasings express the same question with cosmetic changes, counting five identical rows gives that intent artificial influence. The article on how to build an AI visibility prompt bank covers sample design and versioning; here, the unit should already be identified.
A useful unit can be a prompt family with the same user decision, separated by country, language, model and configuration when those conditions change interpretation. For example, “best platform for a small agency” is not automatically equivalent to “best platform for a global enterprise”. If they share a row only for convenience, the weight loses meaning.
Keep at least these fields:
| Field | Control question |
|---|---|
| prompt_id and family_id | Can I tell which row and intent it represents? |
| bank_version | Which question set was active? |
| market, language, model and configuration | Am I mixing non-comparable cells? |
| outcome and extraction rule | What counts as visible, cited or eligible? |
| evidence_date | When was the signal checked? |
| weight_version and weight_reason | Which rule and decision justify the weight? |
If you use one question per family, record that the unit is the prompt. If you use repetitions, summarize the family's outcome with a variability protocol first and then assign the weight. Do not use weighting to repair execution noise.
Combine signals without pretending they are the same thing
Observable demand and business value answer different questions. Search Console may show page exposure in Google results; internal search may reveal what someone already inside your product is trying to find; CRM may show stages, objections or losses; margin may distinguish a profitable sale from a large but unattractive one. No source is the whole market.
Use four separate dimensions:
- D, observable demand or interest: Search Console impressions and clicks when the query is interpretable, aggregated internal searches, coded tickets or interviews and other permitted signals.
- B, business value: potential revenue, margin, retention, loss risk, account size or an explicit commercial consequence.
- S, strategic importance: entry into a priority market, a category the business wants to defend, launch dependency or positioning requirement.
- C, confidence and comparability: source quality, sample size, freshness, classification consistency and comparability across markets.
Normalize raw signals into documented 0-to-100 bands and keep the originals in an internal appendix. For Search Console, record the property, period, filters and whether the query is genuinely interpretable. Impressions are an exposure signal, not a complete demand figure; keep the Search Console performance report documentation alongside the source definitions.
A missing source is not zero. If you have no CRM data for a family, mark B as unavailable and decide how to apply a provisional rule. Putting zero for convenience says it has no value; those are different claims.
Define a simple rule before looking at the ranking
An auditable rule is often better than an opaque model. As a hypothetical starting point, combine the bands like this:
- 35% observable demand, D.
- 35% business value, B.
- 20% strategic importance, S.
- 10% confidence and comparability, C.
The illustrative score is:
score_i = 0.35D_i + 0.35B_i + 0.20S_i + 0.10C_i
Divide by 100 to obtain a relative weight between 0 and 1, or keep the 0-to-100 scale. This formula is not a universal recommendation. It is a hypothesis that the team must be able to defend and that must be dated in weight_version.
Two controls matter. First, cap each source's contribution so one volume signal cannot dominate everything. Second, do not count D and B twice when they come from the same measure. For example, do not count the same pipeline as “commercial-intent searches” and “CRM opportunities” unless you have checked that they add different information.
When confidence is limited, do not increase a row's weight to compensate for intuition. Keep uncertainty as a label and use sensitivity analysis to see whether the decision is robust.
A hypothetical example with six families
The following table does not represent Mentio traffic or customers. It illustrates an internal worksheet in which every dimension has been normalized from 0 to 100 using documented rules.
| Prompt family | D | B | S | C | Illustrative score | Reading |
|---|---|---|---|---|---|---|
| Best platform for a marketing team | 80 | 70 | 60 | 90 | 73.5 | High demand and comparability |
| Comparison with a direct competitor | 62 | 85 | 75 | 85 | 75.0 | Less volume, stronger decision |
| Price and contracting conditions | 40 | 95 | 90 | 80 | 73.3 | Small signal, high impact |
| Risk, security and compliance | 25 | 90 | 95 | 75 | 66.8 | Critical despite low visibility |
| General category definition | 90 | 25 | 20 | 90 | 53.3 | High interest, lower immediate value |
| Setup and support | 35 | 55 | 50 | 70 | 48.5 | Provisional secondary priority |
The 73.3 for price is specific to this worksheet. The example shows that a family with less exposure can remain high when its outcome changes an important decision, and that D, B, S and C can be discussed separately instead of hiding behind one opinion.
Show exclusions as well. If a family has no comparable source in the period, do not leave it in with a zero score without explanation. You can mark it provisional, exclude it from this version's indicator or keep it in a separate coverage report. What matters is that the denominator and the choice are visible.
Turn the score into an indicator with a versioned denominator
For a simple binary rate, define and document the outcome before running the wave. For example, visible_i can be 1 when an answer meets a predeclared mention-and-position condition, and 0 when it does not. Do not change the condition because the first wave looks bad.
One possible weighted indicator is:
weighted_coverage_v = sum(w_i × visible_i × eligible_i) / sum(w_i × eligible_i)
The v matters. The version should identify the bank, model, configuration, extraction rules and weighting. Keep the earlier denominator; do not compare this week's coverage calculated from a different population as if it were improvement.
Show alongside it:
- unweighted coverage;
- weighted coverage;
- number of eligible families and rows;
- sum of eligible weights;
- failed critical prompts;
- excluded rows and reasons;
- period and weight_version.
If you report only the percentage, the reader cannot tell whether it rose because more answers appeared or because difficult rows were removed.
Stop brand and volume from dominating
Branded prompts are often easier to interpret and can accumulate different signals from discovery queries. Non-branded prompts may be closer to category choice. Do not mix them into one total before showing the strata.
A practical structure keeps at least these views:
- branded versus non-branded;
- discovery, comparison, recommendation, price and risk;
- market and language;
- model and configuration;
- critical versus non-critical families.
You can aggregate them with business weights afterward, but the stratum table must remain available. A high total must not hide that the brand fails on price or security questions that block conversion.
The same care applies to bank length. If one family has twenty variants and another has two, do not assume the first is twenty times more important. Weight the family and use variants to estimate stability or coverage, not to inflate representation.
Run a sensitivity test before publishing conclusions
A priority is more reliable when it does not disappear after a small rule change. Build at least two alternatives:
- Base: 35% D, 35% B, 20% S and 10% C.
- Commercial: 20% D, 45% B, 25% S and 10% C.
- Discovery: 45% D, 25% B, 20% S and 10% C.
Compare family order, not only the final percentage. Record position changes and “decision flips”: a family that moves from high to low priority when B changes by ten points needs human review.
Sensitivity analysis checks fragility; it does not replace inferential statistics. If the ranking changes completely, show the owner two possible decisions and explain which assumption separates them. Do not hide instability by averaging the scenarios.
For a small sample, avoid false precision. A family with few observations may score high because of strategic value and provisional confidence; that justifies a test or review, not a claim that the measurement is exact.
Build an evidence card and minimum governance
Every weight change should leave a short record:
- what decision triggered the review;
- which sources were added or removed;
- Search Console period and filters;
- definition of each band;
- formula and coefficients;
- excluded rows;
- approving owner;
- effective date;
- review date;
- expected effect and observed result.
Do not copy CRM personal data into the publication. Work with aggregate categories and keep provenance in the authorized system. If a business source is qualitative, code the rule before looking at the result or label the coding as expert judgment.
A reasonable cadence is to refresh demand signals when new data exists, review weights on a defined monthly or quarterly date and open an exceptional review when the product, market, model or strategy changes. Recalculating every week without changing the version only adds oscillation; changing the rule every week prevents learning.
AI Share of Voice can provide competitive context, but it does not replace this decision record. Before writing the report, preserve the denominator and exclusions so the AI visibility executive report does not turn a score into certainty.
Use weighting to decide, not to hide gaps
A useful output of this method is an explicit work queue:
- Confirm which families are critical.
- Verify missing sources that could change priority.
- Run or repeat eligible rows with the same configuration.
- Review families with ranking flips in sensitivity analysis.
- Create a content, product, PR or research action with an owner.
- Measure again within the same version before changing weights.
The score does not decide what should be published without context. A family with high B and low C may first need better instrumentation. A family with high D and low B may support discovery but not a sales decision. The action depends on the evidence chain.
Weighting also does not replace variability controls. If the answer changes substantially across repetitions, read the family's value alongside its distribution. Review the AI answer variability protocol and preserve the aggregate result without pretending a single run is stable.
Internal publication checklist
Before using the metric in a meeting, confirm:
- [ ] The measurement unit and family are defined.
- [ ] The decision the metric must support is written.
- [ ] D, B, S and C have documented scales and sources.
- [ ] An unavailable value has not silently become zero.
- [ ] The formula, coefficients and weight_version are stored.
- [ ] The denominator contains only eligible rows and keeps its version.
- [ ] Weighted total appears beside the unweighted control.
- [ ] Branded, non-branded, intent, market and model are not mixed without labels.
- [ ] A sensitivity test has been run.
- [ ] Critical families and failures remain visible.
- [ ] Every derived action has an owner and acceptance criterion.
When results must aggregate across brands and products, separate question weights from corporate attribution with this guide to measuring a portfolio in AI.
To compare periods with different demand, retain fixed weights alongside a current view using this guide to AI visibility seasonality and baselines.
FAQ
How much should Search Console weigh against business value?
There is no universal percentage. Start with an explicit scheme, cap each source's contribution and test whether priority changes under alternative weights. The business should be able to explain why a commercial decision can weigh more than an exposure signal.
Do Search Console impressions equal exact demand?
No. They are an exposure signal in the results Google records, not the market's total interest or demand for a question inside an assistant. Use them with internal search, CRM and context, stating the period, filters and limits.
Should branded and non-branded prompts have the same weight?
Not by default. Separate them into families or strata before aggregating. Branded prompts can carry a lot of signal while non-branded prompts may represent discovery; one total can hide that difference.
What if I have no CRM or internal search data?
Record the source as unavailable, not as zero. You can start with a provisional rule based on the evidence you do have, label confidence and set a date to add the missing source. Do not present the provisional weight as historical truth.
Can I change the weights every month?
You can review them, but every change should create a version and preserve the earlier denominator. Compare results within the same version or recalculate the full series under the new scheme; do not mix unrecorded weights.
Does the weighted metric replace mention rate?
No. The weighted metric helps prioritize decisions, while an unweighted rate diagnoses sample coverage. Show both where possible and flag critical prompts separately so an average cannot hide them.
Want to know if AI mentions your brand?
Discover your visibility in ChatGPT, Claude and Gemini in minutes.
Related articles
How to Build a Prompt Bank for Measuring Your Brand's AI Visibility
Build a representative, stable and actionable prompt bank to measure brand mentions, position and competitors across ChatGPT, Gemini and Perplexity.
GEO StrategyAI Share of Voice: How to Calculate It in ChatGPT, Gemini and Perplexity (2026)
How to calculate AI Share of Voice: mentions, position, prompts and intent weighting across ChatGPT, Gemini and Perplexity.
GEO StrategyAI Visibility Executive Report: 2026 Template
Turn AI visibility metrics into a one-page executive report with risks, decisions, owners and next steps for leadership.