GEO Experiments: How to Measure Real Impact (2026)
You update a page, add proprietary data or earn a mention in an industry publication. Two weeks later, your brand appearance rate rises. It is tempting to present that movement as the result of the action. The problem is that the model may have changed, retrieval sources may have shifted and the answers may have fluctuated because of ordinary sampling noise.
A before-and-after chart does not turn correlation into impact. To learn whether a GEO action probably moved visibility, you need to design the experiment before publishing: what should change, for which questions, through which mechanism, against what control and at what threshold you will make a decision.
This protocol begins where the AI brand visibility prompt bank and the AI answer variability benchmark end. The prompt bank fixes the sample; the benchmark estimates noise. The experiment uses both to test one intervention.
Every experiment produces learning that belongs back in the queue: feed it into your AI visibility action backlog to update confidence and reorder priorities.
What a GEO Experiment Can Actually Show
In most cases, you cannot randomly assign users or models to version A and version B. You also do not control the index, web retrieval or provider updates. A GEO experiment is therefore usually a quasi-experiment: it compares a prompt cohort that should respond to the action with another cohort that should not, before and after the intervention.
The result is not perfect causal proof. It is a chain of evidence:
- The selected metric changes in the treated cohort.
- The change exceeds the variability observed in the baseline.
- Control prompts remain stable or move much less.
- The expected mechanism appears, such as a newly cited URL or corrected framing.
- The effect persists across more than one follow-up wave.
The more pieces that align, the more defensible it is to attribute the movement to the action. If only one aggregate percentage improves during a single measurement, you have a signal, not a conclusion.
Start With a Hypothesis That Can Fail
A useful hypothesis forces you to state the action, audience, outcome, window and mechanism. Use this structure:
If we make [specific action] to [asset or source], then [primary metric] will change in [prompt cohort and model] during [window], because [expected mechanism].
Example:
If we publish a verifiable comparison page for agencies and add original data, the citation rate of our URL will increase in tool-evaluation prompts on Perplexity during the four weeks after indexation, because the page will answer that intent directly with retrievable evidence.
The hypothesis can fail. That is what makes it useful. "Improve our AI visibility" does not define an observable result or tell you when to stop a tactic.
Write a rival explanation before running the experiment as well. For example: "the rate may rise because a Perplexity update favors every brand in the category." The control cohort and experiment annotations exist to test that alternative.
Define One Primary Metric and Do Not Change It Later
Choose one primary metric aligned with the action. The guide to GEO metrics for AI visibility provides the catalog; the experiment requires deciding which metric can answer the hypothesis.
| Action | Suitable primary metric | Secondary evidence |
|---|---|---|
| Improve a page for one intent | Citation rate of that URL | Brand mention and cited passage |
| Correct a product fact | Framing accuracy | Supporting source and error persistence |
| Earn coverage in a publication | Citation rate of the domain or URL | Mention, position and competitors |
| Create a comparison page | Share of Voice in comparison prompts | Position and recommendation arguments |
| Strengthen a brand entity | Rate of answers with the correct entity | Name variants and attributes |
You may retain secondary metrics for diagnosis, but do not retrospectively select whichever one increased most. Choosing the outcome after seeing it turns the experiment into a search for coincidences.
Define the analysis unit explicitly, for example prompt_id + run + model + mode + language + market. If you average models or markets before analysis, you can hide the fact that the effect exists only in one environment.
Build a Treated Cohort and Control Prompts
The treated cohort contains the questions whose answers should respond to the action. If you update a page about GEO tools for agencies, include comparison, evaluation and selection prompts for that audience. Do not add broad "what is GEO" questions just to enlarge the sample.
Control prompts belong to the same market, language and model but cover an intent the action does not touch. In this example, they could address measurement for ecommerce or foundational definitions. Their purpose is not to prove that nothing changes; it is to show how much the environment moves without the treatment.
A useful control meets four conditions:
- It runs with the same protocol and schedule.
- It has a similar level of baseline variability.
- It does not share the asset or message being changed.
- It remains relevant to the same brand and category.
Do not call a historical list run in another model, language or cadence a "control group." Those differences prevent you from knowing whether the contrast came from the action or from the design.
The Experiment Brief: Decide Before You Look
Create one experiment brief and freeze it before the intervention.
| Field | Example |
|---|---|
experiment_id |
EXP-2026-07-COMPARISON-01 |
| Problem | The brand is absent from comparisons for agencies |
| Hypothesis | A comparison with original data will increase Perplexity citations |
| Single action | Publish and internally link one comparison URL |
| Treated cohort | 14 selection prompts for agencies |
| Control cohort | 14 prompts for another intent with similar volatility |
| Environments | Perplexity web, EN, United States |
| Primary metric | Citation rate of the treated URL |
| Baseline | Three weekly waves with five repetitions per prompt |
| Follow-up window | Weeks 2, 3 and 4 after confirmed indexation |
| Success criterion | Sustained net improvement and appearance of the expected URL |
| Decision rule | Scale, iterate or stop based on the result |
| Owner | Owner name and review date |
The numbers in the example are operational heuristics, not a universal statistical recipe. A critical or highly variable cohort needs more repetitions. An exploratory test may start smaller as long as it reports its limits.
Step-by-Step Protocol
1. Isolate One Intervention
Decide what changes and what remains fixed. An intervention can update a page, publish a data asset, correct schema, earn an editorial inclusion or resolve an entity inconsistency.
Avoid publishing three articles, changing the product message, launching a PR campaign and restructuring internal links on the same day. If the outcome changes, you will not know which action to keep. When tasks are inseparable, treat them as a declared package and be clear that you are evaluating the package as a whole.
2. Measure a Repeated Baseline
Run treated and control cohorts before the intervention. Preserve full answers, sources, date, model, mode and every available identifier. Repeat each prompt enough times to estimate its normal range instead of relying on one favorable answer.
The baseline should cover more than one point in time. Three weekly waves are usually more informative than fifteen runs in one afternoon because they include temporal variation. If the provider announces a model change during this phase, annotate it or restart the baseline.
3. Publish and Record the Exposure Point
Store the intervened URL or source, its version, publication date and exact changes. When possible, preserve a hash or content snapshot. Confirm that the page responds, is indexable and is not blocked.
Publication time is not always exposure time. A web-connected model cannot retrieve a page it does not know or cannot crawl. Start the follow-up window using the predefined condition, such as confirmed indexation or the first relevant crawl.
4. Freeze the Rest of the Environment
Do not edit the treated asset again during the window unless there is a critical error. Keep the same prompt-bank version, clean sessions, language, market, mode and schedule. Annotate external events: model updates, industry news, competitor launches or access incidents.
Freezing does not mean ignoring the business. It means documenting unavoidable changes so the final analysis does not pretend that a perfect laboratory exists.
5. Run Follow-Up Waves
Repeat the baseline protocol exactly. Do not discard unfavorable answers or replace prompts that stop producing the expected result. If a prompt becomes invalid because the market genuinely changed, mark it and analyze the original set before creating a new version.
Do not stop the experiment at the first positive spike. Complete the planned window or apply a stopping rule written in advance.
6. Calculate the Net Change
A simple reading uses a difference-in-differences calculation:
net change = (treated after - treated before) - (control after - control before)
Consider this illustrative result:
| Cohort | Baseline | Follow-up | Change |
|---|---|---|---|
| Treated prompts | 12% citation rate | 28% citation rate | +16 points |
| Control prompts | 15% citation rate | 19% citation rate | +4 points |
| Estimated net change | +12 points |
This adjustment does not create perfect causality, but it prevents you from attributing a general four-point rise to the action. Report counts and denominators with percentages; a jump from 0 to 50% across two observations is not the same as a stable change.
7. Check the Mechanism
An improvement is more credible when it occurs through the predicted path. If the hypothesis expected a new page to be cited, look for that URL in the answers. If it aimed to correct a fact, inspect the passage and the source supporting the new framing.
When the brand improves but the treated asset never appears and controls rise too, the outcome may be real, but the proposed mechanism is unproven. Keep "visibility improved" separate from "this action caused the improvement."
Interpret Results Without Forcing a Win
Treated Improves, Control Is Stable and the Mechanism Appears
This result is most compatible with the hypothesis. Confirm that it exceeds the baseline range and persists. You can scale the tactic to another intent through a new experiment; do not assume it will work site-wide.
Treated and Control Improve Together
A common factor probably exists: a model change, seasonality, general brand growth or index variation. Calculate the net change and do not give the intervention credit for the full movement.
The Metric Does Not Change, but the Mechanism Appears
AI begins citing the treated URL, but the brand does not gain mentions or position. The action affected the evidence without yet changing the business-facing outcome. Decide whether it needs more time, a clearer proposition or external support.
There Is One Spike and Then a Return to Baseline
Treat it as volatility until it repeats. One positive wave does not justify scaling budget.
Nothing Changes
A null result is useful when the experiment was designed well. The action may not have been retrievable, the mechanism may have been weak, the window may have been too short or the metric may not have responded. Record the lesson and change one variable in the next test.
Minimum Record for an Auditable Result
Keep at least:
- Experiment ID and version.
- Hypothesis and rival explanation.
- Intervened URL, source or asset.
- Exact change and owner.
- Treated and control prompt IDs.
- Model, mode, language, market and session context.
- Full answer, mentions, position, framing and citations for each run.
- Baseline, exposure and follow-up dates.
- Primary metric, denominator and success rule.
- Incidents, external updates and protocol deviations.
- Final decision: scale, iterate, stop or repeat.
This record prevents a promising result from becoming an anecdote that a new team cannot reproduce.
Seven Mistakes That Invalidate Learning
- Changing too many things at once. You get movement but cannot identify the working lever.
- Choosing prompts after the intervention. The sample adapts to the outcome you want to prove.
- Using one run per prompt. You confuse one possible answer with normal behavior.
- Having no controls. Every environmental change looks like credit for the action.
- Mixing models, modes or markets. The average hides opposing effects and breaks comparability.
- Changing the primary metric at the end. Some number will always improve by chance.
- Claiming absolute causality. A quasi-experiment reduces rival explanations; it does not control the entire generative system.
Do not turn this experiment into a complete financial model either. To connect visibility movement with pipeline, revenue and cost, use the GEO ROI and attribution framework after you first validate that the action moved the visibility signal.
From One Experiment to a Learning Program
The value is not one isolated test. It is a body of comparable decisions. Use a shared register to answer:
- Which interventions improve citations, mentions or framing.
- Which models and intents respond.
- How long the first effects take to appear.
- Which outcomes repeat and which were noise.
- Which tactics should be scaled, iterated or stopped.
Every experiment should end with a decision and a next question. If a comparison earns citations but does not improve the recommendation, the next test can work on the argument. If an editorial mention improves one model but not another, investigate each environment's sources before repeating the investment.
Mentio helps teams maintain a stable prompt bank, run recurring measurements across models, preserve answers and compare mentions, position, competitors, framing and sources. That reduces capture work and lets the team focus on the important part: forming a clean hypothesis, controlling changes and making an evidence-based decision.
Pre-Launch Checklist
- The hypothesis states action, cohort, metric, window and mechanism.
- There is one intervention or one declared package.
- The primary metric was selected before results were seen.
- Treated prompts represent one intent.
- Controls share the environment but not the treatment.
- The baseline includes several waves and repetitions.
- Exposure time has a verifiable condition.
- The protocol freezes model, mode, language and market.
- The success rule and window are written down.
- The report will include limitations and rival explanations.
FAQ
What is a GEO experiment?
It is a protocol that defines a hypothesis, a specific action, an affected prompt cohort, control prompts, a baseline, a follow-up window and a success criterion before the intervention. Its purpose is to separate a change compatible with the action from normal noise and external factors.
Is a GEO experiment an A/B test?
It is usually not a randomized A/B test because you do not control which model, index or source each user receives. It is commonly a before-and-after quasi-experiment with treated and control prompts. This is stronger than a simple time comparison, but it does not prove absolute causality.
How many prompts do I need for a GEO experiment?
There is no universal number. The cohort must cover the affected intent without mixing different problems and include enough repetitions to estimate variability. Ten to twenty well-scoped prompts can be a directional starting point, but that is not a statistical guarantee.
How long should I wait after publishing the action?
It depends on web retrieval, crawl frequency and the measured surface. First confirm that the asset is accessible and indexable, predefine the observation window and use several follow-up waves. Models without web retrieval may take much longer or may not reflect the change.
Which prompts should be used as controls?
Use prompts from the same market, language, brand and model that should not be directly affected by the action. They should have similar volatility and value but belong to another intent. If they also change, that helps identify a general update or external factor.
Can a GEO experiment prove that the action caused the change?
It can strengthen attribution when the treated cohort changes, the control stays stable, the expected mechanism appears and the effect repeats. Most GEO experiments are observational or quasi-experimental, so report evidence compatible with the hypothesis rather than absolute causality.
Turn Every GEO Action Into a Measurable Decision
Publishing, earning a mention or correcting a page is not the end of the work. Define what should change first, measure a baseline and compare the outcome with controls. That is how you stop collecting correlations and build a program that learns which actions deserve the next investment.
Once the experiment has produced a result, the AI visibility executive report template helps you turn it into a decision for leadership.
Want to know if AI mentions your brand?
Discover your visibility in ChatGPT, Claude and Gemini in minutes.
Related articles
How to Build a Prompt Bank for Measuring Your Brand's AI Visibility
Build a representative, stable and actionable prompt bank to measure brand mentions, position and competitors across ChatGPT, Gemini and Perplexity.
Practical GuidesHow to Measure AI Answer Variability Without Biasing Your Benchmark
Learn how to repeat prompts, control changes and separate signal from noise to build a reliable benchmark of brand visibility in AI answers.
GEO StrategyHow to Measure the Real ROI of Your GEO Strategy (Beyond Share of Voice)
Attribution, financial KPIs, and a business case to defend your GEO investment to the CFO. ROI calculation model with real 2026 examples.