Back to blogAnalytics screen used while diagnosing a ChatGPT visibility dropGEO Analytics

    How to Diagnose a ChatGPT Visibility Drop

    2026-08-22·17 min read

    The alert is valid: your brand's visibility fell outside its normal range, the run passed quality controls and the incident remains open. Now the hard work starts. Did the sample change? Did ChatGPT change? Did the sources supporting the recommendation disappear? Did your own site stop providing clear evidence? Or did a competitor take the space?

    A useful diagnosis does not pick the most intuitive explanation. It orders hypotheses, looks for tests that can eliminate them and preserves an evidence chain. The goal is not to produce a persuasive story, but to reach a confirmed cause or state precisely what remains unresolved.

    This guide begins where the AI visibility drop alert system ends. It does not redefine thresholds, persistence or severity. It assumes the incident warrants investigation and provides a tree that separates five cause families: measurement, platform, sources, brand and competitors.

    Before the Tree: Freeze the Incident

    Do not investigate a series that changes while you observe it. Freeze the ID, owner and cells; prompt and pipeline versions; model, mode, market and language; repetitions and errors; answers, sources and passages; numerators, denominators and baseline; plus deployments, owned changes and competitors already included.

    That packet is time zero. New runs are added as tests, but they do not replace initial evidence. Any later correction creates a new version while retaining the previous one.

    Order Matters: Eliminate Cheap Causes Before Complex Ones

    Walk through the branches in this order:

    1. Sample and measurement: does the number represent the same question and was it calculated the same way?
    2. Platform or model: is the change broad and specific to ChatGPT?
    3. Sources and retrieval: did the evidence ChatGPT finds or uses change?
    4. Brand: did owned access, content, entity, product or proof change?
    5. Competitors: did another brand gain relatively with an observable explanation?

    The order reduces waste. A broken label can mimic a drop and take minutes to verify; a source or competitor investigation can consume days. Still, do not close a branch on intuition. For each hypothesis, record the test, the expected result if true and the result that would eliminate it.

    Branch 1: Sample, Run and Measurement

    Start by asking whether you are comparing like with like. A drop can appear because new prompts entered the bank, the mix changed, an intent was reclassified, repetitions were missing or the denominator included invalid answers.

    Frozen reproduction test

    Rerun a limited sample of affected cells under the incident conditions. Do not change content or prompts yet. Compare:

    • the complete answer, not just the final label;
    • literal mention, recommendation, position and sentiment separately;
    • retrieved sources and cited passages;
    • previous and current classifier outputs;
    • denominator of valid opportunities;
    • error, blocking and truncated-answer rates.

    Then run the same cohort with the immediately previous pipeline version. If the drop appears only with the new classifier, this is a measurement incident. If it appears only because the bank composition changed, document a series break. If it persists under both versions, the branch remains open.

    Use negative and positive controls

    A negative control is a brand, intent or market that should not be affected. A positive control is a cell where the change should appear if the hypothesis is true. For example, if you suspect a brand-name normalization rule, test one brand with a short name and another with known variants.

    Do not re-estimate normal noise here. That work belongs in the AI answer variability benchmark. Diagnosis uses that profile to decide how many repetitions are needed to reproduce the signal.

    Branch 2: Platform, Model or Mode Change

    If measurement is consistent, look for the fingerprint of an external ChatGPT change. You do not need to know the internal algorithm; you need to test whether the observed pattern fits a platform break better than a change to your brand.

    Build a breadth matrix:

    Cut Signal compatible with platform Signal pointing elsewhere
    Brands Several brands move together Only your brand drops
    Intents Different families are affected Concentrated in one topic or product
    Markets Appears across comparable markets Only one language or country moves
    Surfaces Limited to ChatGPT Other platforms also fall
    Sources Retrieved domain pattern changes Sources remain stable
    Time Simultaneous, persistent break Gradual decline tied to owned events

    Compare ChatGPT with a control platform and unaffected cells. A simultaneous move across many brands only inside ChatGPT makes a platform cause more plausible, but does not confirm it. Record visible versions and require a differential pattern, not a matching date.

    Branch 3: Sources, Citations and Retrieval

    ChatGPT can receive the same question yet ground its answer in different evidence. Compare sources before and after for each cell:

    • domains and URLs present;
    • source type: owned, editorial, review, directory, marketplace or community;
    • passage supporting each claim;
    • page date, freshness and availability;
    • coverage of important attributes;
    • presence of your brand and alternatives in the same source;
    • consistency between source and generated claim.

    Classify every URL as retained, lost, new or changed. A lost source is a clue: test whether it errors, blocks access, changed its text or was simply not retrieved.

    Test availability outside the dashboard

    For owned pages, inspect HTTP status, canonical, indexability, rendered content, robots and logs. The AI crawler access audit explains how to separate declared policy, observed access and delivered content. For external sources, preserve a copy of the public evidence and record any visible title, content, date or structure change.

    The source branch gains strength when the visibility loss is concentrated in questions that depended on a recognizable document set and those pieces of evidence disappear or change. It loses strength when the same sources remain available and support the brand equivalently.

    Branch 4: Brand-Owned Changes

    If the pattern is specific to your brand, inspect changes under your control. Group them so symptoms do not get mixed:

    Technical access

    • errors, redirects, canonicals or retired pages;
    • primary content missing from rendered HTML;
    • robots, CDN, firewall or consent changes;
    • duplicate URLs or inconsistent internationalization;
    • loss of internal links to the page supporting the answer.

    Content and evidence

    • removed prices, availability, specifications or comparisons;
    • vaguer claims or claims without verifiable proof;
    • content that is stale for the question;
    • title, heading or structure changes that dilute the answer;
    • divergence across landing pages, documentation, listings and external profiles.

    Entity and offer

    • changed name, category, product, market or positioning;
    • launch, withdrawal or availability change;
    • inconsistent organization data across sources;
    • lost reviews, recognition or trust signals;
    • new claims that lack independent support.

    Cross-reference each change with the affected cells, but require another test. Temporarily restore a page, correct access on a test URL, supply missing evidence or compare prompts for an unchanged attribute. If the intervention moves the result in the expected direction while controls remain stable, the hypothesis gains strength.

    Branch 5: Relative Competitor Gain

    A brand can retain absolute presence yet lose ground because another takes more recommendations or better positions. It can also fall while the whole market redistributes. Compare within the same cell:

    • your mention and recommendation before and after;
    • each rival's share of voice and position;
    • the competitor that replaced your brand in specific answers;
    • new sources supporting that competitor;
    • changed attributes, price, availability, social proof or coverage;
    • breadth of the gain across intent, market and model.

    The competitive audit in ChatGPT and Perplexity helps investigate a rival in depth. The question here is narrower: does its gain explain this incident?

    Require cell-level coincidence and a mechanism. "The rival appears more" is a description. A causal hypothesis would be: "for purchase questions in Spain, ChatGPT began retrieving two new comparisons that include the rival and omit our certification; introducing equivalent evidence in the control changes the recommendation".

    Build a Falsifiable Hypothesis Matrix

    Do not manage the investigation in a comment thread. Use a shared table:

    Hypothesis Discriminating test If true If false Status
    Prompt mix changed Recalculate with frozen cohort Drop disappears Drop persists Open
    Classifier failed Blind manual labeling Systematic disagreement Agreement holds Open
    ChatGPT-specific change Brand and platform controls Broad break only in ChatGPT Brand-limited pattern Open
    Key source was lost Compare URLs and passages Source disappears in affected cells Still present and equivalent Open
    Owned content changed Reversal or control page Recovery after restoring evidence Result does not change Open
    Competitor advance Cell-by-cell coincidence Rival substitutes with new evidence No consistent substitution Open

    Update status as eliminated, compatible, probable or confirmed. "Compatible" only means the evidence does not contradict the hypothesis. It is not permission to act as if it were true.

    Use Complete Answers as the Evidence Unit

    The aggregate percentage tells you where to look; the complete answer shows what changed. Build a cell-level diff of included brands, order, recommendation strength, attributes, objections, sources, passages and substituted products. When possible, hide which period is considered "good" from the reviewer to reduce confirmation bias.

    Separate Confirmed, Probable and Inconclusive Causes

    Use an explicit standard:

    • Confirmed: a test discriminates the hypothesis, the effect reproduces or appears in independent evidence, and a reversal or intervention behaves as predicted.
    • Probable: several pieces of evidence converge and main alternatives weaken, but reproduction or intervention is missing.
    • Inconclusive: more than one branch still explains the data or available evidence cannot separate them.

    Do not force one reason when there is a chain. A technical change may remove a page; that loss may eliminate a source; the model may then prefer a competitor. Record root cause, contributing factors and impact mechanism separately.

    First 72 Hours Runbook

    • Hours 0-4: freeze the incident, reproduce a valuable subset and inspect denominators and false negatives.
    • Hours 4-24: compare controls, answers and sources; review owned changes and competitive substitution; eliminate incompatible hypotheses.
    • Hours 24-72: test a small reversal or intervention with one control cell, classify the cause and define remediation and recovery.

    Urgency does not justify changing pages, sources and prompts together: the metric may improve while the ability to explain why disappears.

    Choose Remediation After the Cause

    Cause Immediate response Preventive change
    Sample or pipeline Correct data and reissue a versioned series Parity, coverage and classification tests
    Platform or model Adjust expectation and expand controls Model segmentation and break detection
    Sources or retrieval Restore access or replace lost evidence Critical source inventory and change monitoring
    Brand Restore access, clarity or proof Content, entity and deployment QA
    Competitor Address the attribute or evidence driving substitution Cell-level rival and source tracking

    Do not recommend "publish more content" as a universal remedy. If the problem is a broken denominator, a product update or a retired source, that action does not solve the cause.

    Common Mistakes

    • Starting with the most visible rival, which may be a beneficiary rather than the cause.
    • Blaming an update because the date matches, without a differential pattern.
    • Editing several variables during the investigation and losing attribution.
    • Comparing percentages without answers, sources or substitutions.
    • Confusing page accessibility with evidence that it was used.
    • Closing with "algorithm" or declaring recovery after one run.

    Diagnostic Checklist

    • [ ] The alert and original evidence are frozen.
    • [ ] The cohort reproduces under known versions.
    • [ ] Numerators, denominators and classification were reviewed.
    • [ ] Brand, intent, market or platform controls exist.
    • [ ] Complete answers, sources and passages were compared.
    • [ ] Every hypothesis has a test that can eliminate it.
    • [ ] The cause is classified as confirmed, probable or inconclusive.
    • [ ] Remediation, recovery, uncertainty and owners are documented.

    FAQ

    What should I check first after a valid ChatGPT visibility alert?

    Check measurement integrity first: prompt bank version, coverage, repetitions, market, language, model, mode, account, classifier and denominators. If the drop cannot be reproduced under the frozen configuration, you should not investigate content or competitors yet.

    How do I know whether the drop is a sample problem rather than a real change?

    Rerun the affected cohort with the previous and current versions, use a stable control set and compare complete answers rather than only the final percentage. If the change disappears when the sample or classifier is restored, the cause is measurement; if it persists in comparable controls, continue through the cause tree.

    How can I detect whether ChatGPT or the model changed?

    Look for a simultaneous break across several brands and intents under the same conditions, record any observable model or mode change and compare with control platforms. A broad move limited to ChatGPT supports a platform hypothesis, but does not confirm it by itself.

    How do I distinguish source loss from a brand-owned problem?

    Preserve sources and passages before and after. If external domains that supported the brand disappear or lose weight, test the retrieval and source branch. If sources remain available but owned content changed, became inaccessible, lost consistency or no longer supports the claim, test the brand branch.

    When is a competitor gain relevant?

    It is relevant when your loss coincides in the same cells with a repeated rival gain and new evidence can explain the substitution, such as stronger sources, social proof, availability or a better-fit proposition. A competitor rising does not by itself prove that it caused your drop.

    When can I call the cause of a visibility drop confirmed?

    When a test discriminates that hypothesis from alternatives, the effect reproduces or appears in independent evidence, and a reversal or intervention produces the expected result. If there is only temporal coincidence, classify the cause as probable; if several branches remain open, keep it inconclusive.

    Turn an Alert Into a Defensible Explanation

    A drop does not need a fast answer at any price. It needs an investigation that preserves the incident, eliminates measurement failures, compares controls, follows sources and tests the mechanism before intervention.

    Start with the cheapest branch, record what evidence would eliminate each hypothesis and allow the result to remain inconclusive when the data cannot separate alternatives. That is how teams fix real causes instead of reacting to plausible stories.

    Mentio helps retain answers, mentions, positions, sources and competitors by cell so an alert can be investigated with context. Start measuring your AI visibility and turn every incident into a traceable decision.

    Want to know if AI mentions your brand?

    Discover your visibility in ChatGPT, Claude and Gemini in minutes.

    Related articles