Back to blogWooden shelves filled with books from different collectionsGEO Analytics

    AI Source Concentration Risk: How to Measure Dependency

    2026-09-16·10 min read

    Your brand can gain citations while becoming more dependent. When many answers rely on one domain, a syndicated story or a document nobody maintains, aggregate growth hides an important question: which evidence would remain available if that piece disappeared?

    Source concentration risk is not simply having a small number of links. It means relying on a narrow set of evidence for business-critical claims without understanding its quality, continuity or alternatives. An official source may be sufficient for a first-party fact; twenty copies of an unverified claim may not be.

    This guide is for analytics, brand and communications teams that already capture answers and citations. The deliverable is a claim-level dependency matrix, a reproducible concentration calculation and a removal scenario. It is not another list of publishing platforms or a promise to predict the algorithm.

    Separate presence, support and dependency

    A visible URL proves that a link was displayed. It does not automatically establish that the page supports every sentence or supplied the only information used. Before measuring dependency, distinguish a claim-linked citation, an additional link and documentary support that a reviewer actually verified.

    OpenAI's search documentation distinguishes citations from other relevant links in Sources and recommends opening sources to check support and dates. This matters because a long list may include pages that do not substantiate the claim under examination.

    Our guide to AI sources: Reddit, LinkedIn and forums explains the platform landscape. Here the question changes: within your sample, how much observed support depends on each origin, and what remains when it is excluded? Do not attempt to reconstruct training data or every internal model operation.

    Define two counting units first

    For the domain distribution, use one unique incidence per answer, claim and cited domain. If an answer cites the same domain three times for the same claim, count it once. If it contains two distinct claims, keep both pairs and explain that longer answers can contribute more incidences.

    For the loss scenario, use the answer-claim pair as the unit. Include it in the initially supported set only when a reviewer has verified at least one source that substantiates the claim. Record contradicted, unsupported and unreviewed claims separately.

    Freeze the question bank, market, language, surface, mode, date and repetitions. Separate planned executions, valid answers, answers with citations and reviewed pairs. A valid uncited answer is not a failed capture; neither justifies inventing an unknown domain to complete a denominator.

    The URL record can start with a Perplexity citation tracking workflow. This protocol adds claim-level support review and relationships between sources, rather than counting brand mentions again.

    Preserve four dependency maps

    Do not compress all provenance into a domain column. One platform may host independent authors; multiple domains may share an owner; several articles may depend on a single study.

    Map What it groups Caution
    URL and domain Visible document and website Keep original and normalized URLs; do not merge pages with different content
    Publisher or owner Documented editorial control Shared ownership does not prove identical editorial decisions
    Story or origin Republication and shared primary evidence Document the connection; similar wording is insufficient
    Source type First-party documentation, media, community or directory Use a stable classification and preserve unknown cases

    Retain hostname, registrable domain and, where necessary, author or publication space. A shared hosting domain does not identify one publisher. Do not combine unknown owners into a fictional organization: report the unresolved share and explicit scenarios.

    Syndication needs an origin label, not a silent correction. Preserve the distribution of domains actually displayed and calculate a separate distribution of deduplicated stories. When changing grouping, deduplicate incidences again within each answer-claim-group; do not add overlapping domain shares as though they were independent units.

    Calculate Top 1, Top 3 and HHI without inventing a traffic light

    Within a comparable measurement cell, divide each domain's incidences by all eligible incidences. Top 1 is the largest share; Top 3 is the sum of the three largest. Report the incidence total and number of domains too.

    For the full distribution, HHI can be calculated as the sum of squared p_i shares, expressed between 0 and 1. This descriptively adapts the HHI formula documented by the DOJ; it is not a GEO risk standard. Do not import competition-law thresholds into AI citation analysis.

    On this scale, a single source produces 1 and four equally weighted sources produce 0.25. With no incidences, the result is unavailable. A lower HHI does not guarantee accurate, current or independent sources.

    Fictional example: one cell contains 80 valid answers, 60 with eligible citations and 100 deduplicated incidences. These are not Mentio results or an industry benchmark.

    Fictional domain Incidences Share Squared share
    A 60 0.60 0.36
    B 20 0.20 0.04
    C 10 0.10 0.01
    D 10 0.10 0.01
    Total 100 1.00 0.42

    Top 1 is 60%, Top 3 is 90% and HHI is 0.42. Coverage of answers with citations is 60/80, or 75%. A's share does not mean it appears in 60% of all answers or that losing it would cause a 60% decline: the denominator is incidences, not answers, users or sales.

    Segment before interpreting the average

    Calculate by surface and mode, language, market, question family and claim. A balanced aggregate can conceal one model's near-total reliance on a directory while another uses official documentation.

    Also show concentration by source type and verified origin. Each view answers a different question. A high first-party documentation share may be reasonable for commercial terms; the same dependency deserves more scrutiny for a comparative claim requiring external evidence.

    Compare periods using the same protocol and a common panel. Report windows separately when the AI product, sample or capture coverage changes. Repetition helps assess stability, but answers to one prompt do not become independent users and cannot justify a universal threshold.

    Simulate removal without predicting model behavior

    Freeze the set of initially supported answer-claim pairs. For each scenario, exclude evidence from the selected domain, publisher or origin from the record. Count how many pairs retain at least one sufficient verified source and how many lose all observed support.

    Documentary exposure = pairs losing all verified support / initially supported pairs. Keep the original denominator. Do not report improvement merely because unsupported cases were removed from the report.

    Another fictional example, independent of the previous one: 40 pairs are initially supported. Review confirms that every useful B source copies one A story. The sets without an alternative when removing A and B separately are disjoint; four additional pairs have support only from A and B together.

    Scenario Pairs losing all support Pairs retaining support Exposure
    Remove domain A only 12 28 30%
    Remove domain B only 2 38 5%
    Remove common origin A + B 18 22 45%

    The third scenario is not the sum of the first two: it includes four pairs whose apparent alternative was another copy of the same story. The actual calculation must use sets of pairs, not sums of percentages.

    This test does not remove pages from the internet or establish what a model would retrieve afterward. AI could find new sources, answer without citations or change its answer. To observe behavior, run a separate controlled wave and retain its results independently from the documentary scenario.

    Prioritize claims, not a lower number

    Read exposure alongside criticality, freshness, support quality and ability to intervene. You do not need an invented composite score to distinguish a minor specification from a central commercial claim with no verifiable alternative.

    For prices and terms, strengthen the official source and change control. For a comparison, check methodology and independent evidence. When multiple pages copy an incorrect figure, correct the origin and record republications; adding domains does not repair the evidence.

    If legitimate external corroboration is needed, use the Digital PR for AI visibility process. Do not pursue links merely to lower Top 1: you can diversify URLs while retaining complete dependence on the same underlying fact.

    If a loss has already occurred, move to diagnosing a ChatGPT visibility drop. This article identifies prior exposure; it does not replace causal incident investigation.

    Deliver an actionable, reviewable record

    Each risk should include claim, surface, period, critical source, grouping evidence, scenario, numerator, denominator, limitation, owner and review date. Attach affected pairs and the passage that verifies support while respecting source rights.

    A useful action might be to review a compatibility claim whose only current evidence is an external document, confirm a maintained primary reference and repeat the support check afterward. Closing the risk requires reviewed evidence, not merely a lower domain share.

    If you use Mentio to observe visibility, keep these checks alongside the analysis. This guide does not announce automatic HHI calculation, owner grouping or a source-removal simulator in the product. Verify which fields your tool exports and complete any missing review.

    The objective is not to make every source equally important. It is to understand which material claims depend on fragile evidence and which specific decision reduces that fragility without sacrificing accuracy.

    FAQ

    What is source concentration in AI answers?

    It is the extent to which observed citations accumulate in a small number of domains, publishing groups or origins. It describes a specific sample; it does not reveal every internal model source or prove fragility on its own.

    Does a high HHI mean I will lose visibility?

    No. HHI summarizes a citation distribution, not the probability of a decline. Interpret it alongside claim criticality, support quality and removal scenarios, without importing legal competition thresholds.

    Do multiple URLs mean independent sources?

    Not necessarily. They may share a publisher, republish one story or rely on the same original data. Keep URL, domain, owner and origin separate, and document cases where independence cannot be verified.

    How should I handle uncited answers or failed captures?

    Record a valid answer without citations as an observed state, separately from capture failures. Calculate concentration only from eligible citations and disclose sample coverage; with no citations, HHI is unavailable, not zero.

    What does a source-removal test measure?

    It measures which answer-claim pairs would lose all verified support when a source is excluded from a frozen record. It is a documentary exposure scenario, not a prediction of AI behavior or a causal estimate of traffic loss.

    Should I always reduce the share of my own website?

    No. Official documentation may be the right source for prices, terms or specifications. Prioritize accuracy, availability and change control; seek independent corroboration when the nature of the claim actually requires it.

    Want to know if AI mentions your brand?

    Discover your visibility in ChatGPT, Claude and Gemini in minutes.

    Related articles