How to Monitor a Product Launch in ChatGPT, Gemini and Perplexity
On launch day, your website, announcement, partner pages, comparisons and market conversations begin telling a new story. ChatGPT, Gemini and Perplexity may pick up parts of that story at different speeds. They may also retain the old name, mix two versions, or claim that the product is not yet available.
One isolated check cannot tell you whether the launch is being understood. You need to compare what models said before publication, what changes during launch and what becomes stable afterward.
This guide turns monitoring into a time-bound operation. It starts with an AI visibility tracker that is already configured and adds what a launch requires: a dated fact sheet, risk prompts, a more intensive cadence, alerts and owners.
Why a launch is not ordinary tracking
Steady-state measurement looks for weekly or monthly trends. A launch is a known event that changes many sources at once and concentrates risk into a few days.
The team needs to answer four questions:
- Discovery: Does the product appear when someone asks about the problem or category?
- Understanding: Does AI explain what it is, who it serves and how it differs?
- Freshness: Does it recognize current availability, name, version and facts?
- Risk: Does it invent, merge or retain facts that could damage a decision?
You are not demanding that models update instantly. You are observing the lag, locating where the story fails and deciding which evidence or source needs attention.
Define launch truth before opening the dashboard
Do not score answers against what each person remembers from the campaign. Create a launch truth sheet approved by product, marketing and, when necessary, legal.
| Field | Example | Risk if wrong |
|---|---|---|
| Canonical name | Mentio Signal | Confusion with the parent brand |
| Category | AI visibility measurement platform | Incorrect classification |
| Audience | B2B marketing teams | Recommendation to the wrong audience |
| Problem | Detect what AI assistants say about a brand | Distorted promise |
| Differentiator | Recurring comparison across models and competitors | Generic message |
| Availability | Spain and Mexico from July 29 | Offer shown in an unsupported market |
| Price or plan | Only if public and stable | Incorrect commercial decision |
| Previous-product relationship | Replaces X; complements Y | Version confusion |
| Approved sources | Product page, documentation and announcement | Fragmented evidence |
| Owner | Product Marketing | Contradictions go unresolved |
The sheet does not replace an AI brand entity audit. It uses that audit as its foundation and limits evaluation to facts that change with this launch.
Assign each fact:
- high, medium or low criticality;
- a primary URL;
- an effective date;
- markets and languages;
- acceptable wording;
- phrases that indicate an error or confusion.
Build a launch-specific prompt matrix
Do not add the product name to twenty prompts and call it monitoring. Combine branded and unbranded questions to observe discovery and understanding.
| Block | What it tests | Example |
|---|---|---|
| Unbranded problem | Discovery | How can I learn what AI assistants say about my brand? |
| Category | Classification | Which platforms measure visibility in ChatGPT, Gemini and Perplexity? |
| Comparison | Differentiation | Compare the new product with two real alternatives |
| Branded definition | Understanding | What is [product] and who is it for? |
| Availability | Freshness | Is [product] available in Mexico? |
| Compatibility | Fit | Does it integrate with [relevant system or workflow]? |
| Migration | Version relationship | Should I keep using [previous product]? |
| Price or plan | Commercial risk | How much does [product] cost? |
| Identity confusion | Disambiguation | Does [product] belong to [company]? |
| Control | Environmental stability | A related question that should not change |
Start from a versioned AI visibility prompt bank, then create a launch cohort. Label intent, market, language, model, criticality and expected fact. Freeze wording and configuration before T-7 so a prompt edit cannot look like adoption.
Phase 1: record the baseline from T-14 to T-4
The baseline describes what models knew before public exposure. Run the cohort under comparable conditions and retain full answers.
Record for every run:
- date and time;
- prompt and version;
- model, mode and web access;
- country and language;
- complete answer;
- brand and product mentioned;
- position or prominence;
- cited sources and URLs;
- narrative classification;
- correct, missing or wrong facts;
- auditable capture or reference.
Repeat every condition several times. Generative systems vary, so apply the AI answer variability protocol within each model and market. A baseline is not one answer collected the night before.
This phase should separate three states:
- The product does not yet appear, as expected.
- It appears because public information already exists.
- It appears with incorrect or prematurely released facts.
The third state needs immediate review. Do not wait until T0 to discover an indexed page with the wrong name, price or availability.
Phase 2: check readiness from T-3 to T-1
Before launch, validate that approved sources are publishable, consistent and accessible. You are not measuring success yet; you are checking that the evidence system is not broken from the start.
Review:
- one canonical URL per product and market;
- consistent name, category and description;
- dates and availability without contradictions;
- aligned documentation and FAQ;
- valid structured data where appropriate;
- links from the parent brand and related products;
- partners or distributors using the correct version;
- no old page that still looks current.
Run a short wave at T-1 with the baseline configuration. Save any pre-announcement change. If you change prompts, sources or conditions, create a new version and do not mix its results into the baseline.
Phase 3: monitor the T0 to T+3 window
Record the public launch time. From that point, run dated waves rather than improvised checks by different team members.
A practical cadence:
| Time | Objective |
|---|---|
| T0 | Capture the state at publication |
| T+6 hours | Detect early critical errors |
| T+1 day | Compare initial adoption by model |
| T+3 days | See whether the signal persists or was a spike |
Not every team needs a run every six hours. Reserve high frequency for critical facts, priority markets or campaigns with substantial exposure. Keep enough repetitions to avoid mistaking one answer for a change.
Calculate coverage before interpreting each wave:
coverage = valid runs / planned runs
If a model or configuration fails, mark the wave incomplete. Do not present an aggregate improvement built from fewer observations.
Classify adoption, absence and error
A binary mention does not explain whether the launch was understood. Classify each answer with a shared taxonomy:
| Status | Criterion |
|---|---|
| Adopted and correct | Name, category, audience and main fact match |
| Correct but incomplete | No errors, but a relevant element is missing |
| Stale | Keeps an old version, price, date or availability |
| Confused | Mixes the product, company, version or competitor |
| Incorrect | States a fact that conflicts with the approved sheet |
| Mentioned without context | Names the product but does not explain its fit |
| Absent | Does not appear in a question where it could be relevant |
| Not evaluable | Empty answer, blocked run or invalid configuration |
Separate mention rate from narrative adoption rate:
adoption rate = adopted and correct answers / valid answers
A brand can gain mentions while accuracy gets worse. Show critical-error volume and coverage next to every rate.
Measure the delta against baseline
For every wave and condition, compare:
- change in mention rate;
- change in narrative adoption;
- change in position or prominence;
- number of correct facts;
- new and resolved errors;
- newly cited sources;
- confusion with previous products;
- dispersion across repetitions.
Do not merge models and markets automatically. A launch may be well understood by Perplexity in Spain and poorly described by ChatGPT in Mexico. Preserve the condition-level view and aggregate only with explicit business weights.
The delta shows a temporal association, not causality. If you need to attribute a change to a specific intervention, design a GEO experiment with controls and an expected mechanism.
Define alerts before seeing results
Escalation rules should exist before an uncomfortable answer appears.
| Level | Example | Action |
|---|---|---|
| Critical | False price, availability, safety or identity | Open an incident immediately and assign an owner |
| High | Repeated confusion in two models or a priority market | Review sources the same day |
| Medium | Incomplete message repeated across two waves | Improve evidence and observe the next wave |
| Low | One absence or position variation | Record without reacting |
One possible rule:
- Escalate any critical error even if it appears once.
- Escalate a high-severity error if it repeats in the same condition or appears in two models.
- Do not escalate one absence without reviewing repetition, relevance and coverage.
- Close an incident only after the source is corrected and two comparable waves no longer reproduce the problem.
Correction may require updating a primary page, aligning documentation, removing an old URL or contacting an external source. Use the broader guide to fixing wrong AI information about a company for the general process; here, every action must remain tied to a launch incident.
Phase 4: confirm stabilization from T+4 to T+14
Reduce frequency and retain comparable checkpoints at T+7 and T+14. Look for stability, not a perfect score.
The launch can close when:
- Planned coverage is complete.
- Primary metrics do not change materially across two waves.
- No critical errors remain open.
- The correct narrative appears consistently in priority conditions.
- Dominant sources have been identified.
- Every incident has a resolution, owner or explicit risk acceptance.
Save the closing wave as the new operating baseline. Retire prompts tied only to a date or temporary availability; category, comparison and product-relationship prompts move into recurring tracking under a new version.
Example: launching a module for a B2B SaaS platform
A company launches an analytics module for customers of its main platform.
At T-7, the new name is absent, but ChatGPT assigns its function to the old product. At T0, Perplexity cites the announcement and describes the function correctly, while Gemini keeps the previous architecture. At T+1, two answers claim that the module is sold separately even though it is available only within a specific plan.
The team does not conclude that "Gemini is worse" or rewrite the entire campaign. It opens a commercial incident, discovers ambiguous wording on the pricing page, corrects it and aligns the documentation. At T+3, it repeats the same cohort. The error disappears in Perplexity and ChatGPT but remains in one Gemini repetition, so it stays under observation until T+7.
The system's value is not declaring a winner. It connects a specific answer to a fact, source, owner and next check.
Assign owners and keep one incident log
| Role | Responsibility |
|---|---|
| Product Marketing | Truth sheet and narrative |
| Product | Features, compatibility and versions |
| SEO/GEO | Cohort, runs, sources and links |
| PR/Comms | Announcement and external sources |
| Legal/Compliance | Regulated or sensitive facts |
| Analytics | QA, deltas and coverage |
| Incident owner | Correction and closure |
Use one log with:
incident_id, date, prompt, model, market, answer, affected_fact, severity, likely_source, owner, action, status, next_measurement
The leadership summary should contain changes, risks and decisions rather than hundreds of screenshots. Move the closeout into an AI visibility executive report and keep the dataset as an auditable appendix.
Mistakes that invalidate launch monitoring
- Starting on launch day without a baseline.
- Changing prompts between waves without creating a version.
- Running only one answer per condition.
- Treating any mention as correct narrative adoption.
- Merging countries, languages and models into one percentage.
- Ignoring incomplete coverage or access failures.
- Reacting to every absence as a crisis.
- Failing to separate the new product, old version and parent brand.
- Correcting sources without recording the incident that caused the change.
- Closing monitoring without turning T+14 into the new baseline.
Operational checklist
- [ ] The truth sheet is approved and dated.
- [ ] Every fact has a source, criticality and owner.
- [ ] The cohort covers discovery, understanding, freshness and risk.
- [ ] Prompts and configuration are frozen before T-7.
- [ ] There is a repeated baseline by model, market and language.
- [ ] T0, T+1, T+3, T+7 and T+14 are scheduled.
- [ ] The taxonomy separates adoption, absence, staleness, confusion and error.
- [ ] Alerts are defined before results are reviewed.
- [ ] Every incident retains the answer and evidence.
- [ ] Closeout creates a new baseline and an open-issues list.
FAQ
When should I start monitoring an AI product launch?
Start seven to fourteen days before the public date. That window lets you record a baseline, check for premature or incorrect information, and freeze the sample before the campaign, press coverage and product pages change the environment.
Which prompts should launch monitoring include?
Include category, problem, comparison, compatibility, availability, price, migration and previous-product relationship prompts. Add controls that should not change and risk prompts about facts that would cause harm if AI explained them incorrectly.
How often should I run prompts during the launch?
Use a more intensive cadence from T0 to T+3, with multiple repetitions per condition, then reduce frequency as answers stabilize. Keep comparable checkpoints at T-7, T-1, T0, T+1, T+3, T+7 and T+14.
How do I know whether AI has adopted the launch narrative?
Adoption requires more than mentioning the name. The answer should correctly connect the product, category, audience, availability and main proposition without confusing it with a previous version. Measure that match against an approved fact sheet.
Should one incorrect answer trigger an alert?
Record every incorrect answer, but not every one requires an incident. Escalate critical price, availability, safety or identity errors immediately; for lower-severity errors, require repetition in the same condition or occurrence in more than one model.
When can the launch be considered stable?
When planned coverage is complete, primary signals stop changing materially between waves, no critical errors remain open, and the fact sheet appears consistently in priority conditions. Document closure and keep the dataset as the new baseline.
Turn the launch into a measurable window
A launch does not end when the announcement goes live. It ends when the market can find and understand the offer without distorting its critical facts. Record the before state, observe the change and keep enough evidence to decide what to fix, what to wait for and what to add to recurring tracking.
Want to know if AI mentions your brand?
Discover your visibility in ChatGPT, Claude and Gemini in minutes.
Related articles
AI Brand Entity Audit: Names, Products, Aliases and Facts
Build a verifiable map of brand names, products, aliases and facts to find ambiguity before optimizing your brand's visibility in AI answers.
Practical GuidesHow to Measure AI Answer Variability Without Biasing Your Benchmark
Learn how to repeat prompts, control changes and separate signal from noise to build a reliable benchmark of brand visibility in AI answers.
Practical GuidesAI Visibility Tracker: What a Serious Tool Should Measure (2026)
An AI visibility tracker should not just count mentions: it should measure position, framing, competitors, sources and changes by model. A practical guide.