How to Track AI Product Recommendations at SKU Level
Your brand appears in a ChatGPT answer. The result looks positive until you open the catalog: the AI recommends a retired model, assigns a feature to the wrong SKU and omits the product that is actually available in the requested market. Brand visibility cannot reveal that problem.
Product measurement needs a more precise unit. Every observation should connect a real prompt to an identifiable SKU or variant, a position, recommendation strength, the attributes used to justify it, availability state and the substitutes that took its place. Without that structure, a mention average mixes products, countries and intents that do not compete in the same decision.
Google Search Console showed Mentio 29 impressions for the exact query chatgpt monitoring for product pages and 19 for how to make products visible in gemini in the snapshot recorded on August 16, 2026. These are small, separate signals, not a market estimate. They support one operational task: moving from knowing whether a store is visible to testing which products each system recommends and under what conditions.
This guide starts with an existing catalog. The general strategy for making a store eligible belongs in GEO for ecommerce; here we will build the measurement system that reveals what happens SKU by SKU.
Why a Brand Tracker Is Not Enough
A brand can improve its mention rate while the catalog that matters most gets worse. The AI may cite the company in informational prompts while omitting high-margin SKUs in purchase decisions. It may also recommend the correct family but a variant that is unavailable in the requested country.
Separate at least these levels:
| Level | Question it answers | Risk when mixed |
|---|---|---|
| Brand | Does the company appear? | Hides which product generated the mention |
| Family | Does the line or model appear? | May ignore size, capacity or generation |
| SKU | Does the specific sellable unit appear? | Requires alias and variant resolution |
| Offer | Can it be bought now in that market? | Price, stock and seller change quickly |
The right level depends on intent. “Best headphones for travel” may be evaluated by model. “Black noise-cancelling headphones delivered tomorrow under $200” requires a compatible variant and offer.
Define the Decision Before Measurement
Do not begin by loading the full catalog. Write down the decision the monitoring system must support. For example:
- detect missing SKUs in high-value prompts;
- identify incorrect attributes that change selection;
- test whether an out-of-stock product remains recommended;
- measure substitution between owned and competitor products;
- compare category coverage across ChatGPT, Gemini and Perplexity;
- monitor differences by country, language or availability.
Every decision needs a metric, cadence and owner. If the team does not know what it will do when a SKU disappears, it does not yet have a monitoring system; it has a collection of screenshots.
Recommendation monitoring observes what models answer. For Shopify store configuration and sales attribution, use the Shopify Agentic Storefronts runbook.
Freeze a Catalog Snapshot
The series is comparable only if you know which products could compete on each date. Create a versioned snapshot containing:
- stable internal identifier and commercial SKU;
- parent product and variant;
- name, aliases and generation;
- category and subcategory;
- attributes that condition selection;
- price, currency and range;
- stock and eligible markets;
- canonical product URL;
- launch, retirement or replacement date;
- source and time of the latest update.
Do not use the visible title as the key. It may change for SEO, translation or merchandising. The identifier must survive those edits and map the language the AI uses back to the product.
Separate Identity From Offer
The product and its offer are not the same thing. A model may exist but be out of stock, sold through several retailers, or have different prices and delivery by market. Keep a relatively stable identity table and a dated offer table. This separates “the AI does not know the SKU” from “it knows it but recommends an offer that is no longer valid”.
Build Comparable Cells
The recommended observation unit is:
SKU or family × intent × model × market × language × date × repetition
Do not drop dimensions just to make the dashboard cleaner. A recommendation can change between Spain and Mexico, Spanish and English, or a general prompt and a price-constrained prompt.
Define before execution:
- Eligible catalog: SKUs that could answer the intent on that date.
- Competitive set: comparable owned and external alternatives.
- Configuration: model, mode, account, location and any controlled personalization.
- Denominator: valid answers, not total attempts when failures occurred.
- Cadence: frequency compatible with catalog volatility and the decision.
When a cell changes, open a new version. Do not splice a new market, taxonomy or model into the same series without marking a break.
Design Prompts Around Shopping Jobs
A useful prompt bank does not repeat the product name. It represents decisions in which a SKU can be selected or rejected. Include families such as:
| Job | Example | What it reveals |
|---|---|---|
| Discovery | “Best compact coffee makers for a small apartment” | Initial inclusion and category |
| Comparison | “Model A versus Model B for daily use” | Substitution and decisive attributes |
| Constraint | “Laptop under $1,200 with 32 GB” | Price and specification accuracy |
| Compatibility | “Accessory compatible with…” | Product relationships |
| Availability | “What can I buy today in Spain?” | Stock, market and offer |
| Risk | “Which option should I avoid if I need…” | Objections and disqualification |
Assign every prompt to an intent and eligible set before seeing the answer. Do not change criteria afterward to turn an unexpected mention into a success.
Turn the Answer Into a Structured Observation
Store the complete answer and extract separate fields:
- mentioned SKU or family;
- order of appearance;
- primary recommendation, alternative, comparison or rejection;
- endorsement strength;
- positive and negative attributes;
- quoted price, stock, market and seller;
- URL or source when available;
- substitute that occupied the expected slot;
- uncertainty or conditional language;
- factual error and severity.
Do not reduce everything to “mentioned yes/no”. A SKU can appear only as an option to avoid. Treat inclusion, position and recommendation as separate variables.
Resolve Names Without Inventing Matches
The AI may abbreviate a model, omit its generation or merge two variants. Use an alias dictionary with three outcomes:
- exact match: the answer identifies the SKU or variant;
- family match: it identifies the line but not the sellable unit;
- ambiguous: there is insufficient evidence to assign a SKU.
Ambiguous cases should not be assigned automatically to the most popular product. Route them to review and preserve the text that caused the uncertainty.
Measure Inclusion, Position and Strength Separately
Three basic metrics answer different questions:
inclusion rate = valid answers including the SKU / eligible valid answers
recommendation rate = valid answers recommending the SKU / eligible valid answers
conditional average position = sum of positions / answers where it appears
Conditional average position does not represent absences. Publish it with coverage and inclusion. A product appearing once at position 1 does not automatically outperform one appearing in eight of ten answers at position 3.
For recommendation strength, use a defined scale such as primary, alternative, neutral mention and advised against. Document edge cases and review a blind sample to test whether classifier and reviewers apply the same rule.
Audit Attributes, Not Only Names
The AI can recommend the correct SKU for the wrong reason. Build a critical attribute matrix by category:
- technical specification;
- use case or audience;
- price and currency;
- size, color or capacity;
- compatibility;
- availability and delivery;
- warranty, certification or policy;
- relevant limitation.
Classify every claim as correct, incorrect, stale, unverifiable or omitted when the decision required it. A color error does not have the same impact as nonexistent compatibility; add severity based on commercial risk, returns, safety and trust.
Separate Visibility From Purchasability
A SKU can be visible without being a purchasable option. Use a compound state:
| State | Meaning | Analysis action |
|---|---|---|
| Visible and purchasable | Valid identity and offer | Monitor position and rationale |
| Visible, uncertain offer | Correct product, unverifiable price or stock | Review freshness and source |
| Visible, ineligible | Out of stock, retired or outside market | Measure risk and substitution |
| Not visible, eligible | Could solve the intent but was omitted | Investigate coverage and evidence |
| Ambiguous | Family or variant cannot be resolved | Human review |
The Perplexity Shopping guide explains how to prepare a catalog for that surface. This system does not assume a shopping card: it records conversational recommendations across platforms and separately tests whether the observed offer is still valid.
Record Substitutes and Cannibalization
When a SKU is absent, ask what took its place. Classify the substitute as:
- another variant of the same product;
- an owned product from another family;
- a direct competitor;
- an alternative from another category;
- a recommendation without a specific product;
- no recommendation.
Owned substitution may be positive for revenue and negative for margin, inventory or launch strategy. Do not count it automatically as catalog success. Preserve the expected SKU, observed SKU and attribute that appears to explain the change.
Aggregate Without Hiding the Long Tail
A sales-weighted average can let five best sellers hide hundreds of SKUs that are never recommended. Publish at least:
- eligible, measured and valid-data SKUs;
- percentage with zero inclusions;
- median and percentiles of recommendation rate;
- distribution by category and intent;
- weighted and unweighted results;
- variant coverage;
- attribute errors by severity;
- owned versus competitor substitution.
If you apply revenue, margin or strategic-priority weights, freeze the version and show the effect of weighting. The GEO metrics guide helps separate presence, position and sentiment; here those metrics operate on comparable catalog units.
Separate Real Change From Variability
AI answers vary even when the catalog does not. Before declaring that a SKU gained or lost visibility:
- repeat prompts under the same configuration;
- preserve the same eligible set;
- compare equivalent windows;
- review failures and truncated answers;
- require persistence for small changes;
- test whether the pattern appears across several shopping jobs.
Calibrate repetitions and limits with the AI answer variability benchmark. If the catalog, model or prompt mix changes, mark a series break instead of presenting false continuity.
Recommended Operating Cadence
A steady-state operation freezes catalog, offers, prompts and configuration before execution; records coverage, failures and versions; resolves aliases and validates attributes; compares only equivalent cells; separates absence, error, invalid offer and substitution; assigns an owner; and closes after repeating a sample with recovery evidence.
Do not use this cadence for a known dated event. A launch requires a baseline and T-14 to T+14 windows; that work is covered in AI product launch monitoring.
Define Catalog-Specific Alerts
Open a candidate when a priority SKU loses inclusion, a retired variant reappears, a critical error increases, a rival persistently replaces the SKU or valid coverage drops.
Thresholds should combine coverage, magnitude and persistence. The general AI visibility drop alert system explains how to separate noise from an incident. For a catalog, severity should also account for sales, margin, stock, returns and misinformation risk.
Common Mistakes
- Counting a brand mention as a recommendation for all its products.
- Assigning an ambiguous family to the best-selling SKU.
- Mixing positions and absences in one average.
- Ignoring that the product was out of stock or outside the market.
- Changing catalog, prompts and classifier within the same series.
- Measuring only best sellers and presenting the result as total coverage.
- Treating every appearance as positive endorsement.
- Comparing platforms without controlling market, language and intent.
- Editing a page and attributing the next change without controls or repetitions.
Implementation Checklist
- [ ] A stable ID exists for family, variant and offer.
- [ ] The catalog snapshot has a date, market and source.
- [ ] Every prompt is assigned to an intent and eligible set.
- [ ] Cells retain model, language, market, date and repetition.
- [ ] Inclusion, position, strength and purchasability are separate variables.
- [ ] Ambiguous aliases go to review.
- [ ] Critical attributes are validated against a dated source.
- [ ] Owned and competitor substitutes are recorded.
- [ ] Aggregates show coverage, distribution and zeros.
- [ ] Changes compare only equivalent versions.
- [ ] Every alert preserves the complete answer and an owner.
FAQ
What is the minimum unit for measuring AI product recommendations?
The minimum unit should be a versioned SKU observation by intent, model, market, language and date. It should preserve the prompt, complete answer, position, mentioned attributes, observed availability and any substitute so the result can be reproduced and compared.
How many prompts and repetitions do I need per SKU?
There is no universal number. Start with intents that represent discovery, comparison, constraint and decision, then calibrate repetitions against the observed variability of each model. Higher-value or higher-risk SKUs need more coverage, but every comparison must preserve the same cell and a valid denominator.
Should I count a recommendation if the product appears without availability?
Record the inclusion, but do not mix it with a purchasable recommendation. Use separate states for included, recommended, available, verifiable price and correct link. An out-of-stock SKU can retain informational visibility while still failing purchase intent.
How should I handle colors, sizes and other variants?
Keep one family identifier and another variant identifier. Aggregate to the parent only when the prompt does not require a specific variant; if it asks for size, color, capacity or market, evaluate the child SKU. Document equivalence rules so a family does not receive credit when the requested variant does not exist.
Is the first mention always the best recommendation?
No. Position must be interpreted with recommendation strength, attributes, warnings and intent. A product cited first as an option to avoid is not a primary recommendation. Preserve the complete answer and classify inclusion, order and endorsement separately.
How do I aggregate many SKUs without hiding the long tail?
Report distribution and coverage, not only an average. Separate eligible and measured SKUs, show median and percentiles, disclose the share with zero recommendations, and break results down by category, intent, market and model. If you apply commercial weights, retain an unweighted view and version the weights.
SKU-level measurement observes the outcome; the ChatGPT product feed covers the delivery that precedes it. Connecting both supports investigation without claiming that an accepted row caused a recommendation.
Turn the Catalog Into an Explainable Series
Product visibility is not a list of names that appeared once. It is a versioned series connecting eligible catalog, intent, answer, recommendation, attributes, offer and substitution.
Start with one category and one concrete decision. Freeze the SKUs, retain complete answers, separate inclusion from purchasability and aggregate only after reviewing the long tail. That is how you discover which product lost an opportunity, what replaced it and what evidence the team needs to act.
Mentio helps organize mentions, positions, answers and competitors by prompt and model. Start measuring your AI visibility and build catalog monitoring that reaches the SKU.
Want to know if AI mentions your brand?
Discover your visibility in ChatGPT, Claude and Gemini in minutes.
Related articles
GEO for Ecommerce: Get Your Store Recommended by AI
When shoppers ask ChatGPT where to buy what you sell, do you appear? GEO guide for ecommerce to get recommended in ChatGPT, Gemini and Perplexity.
GEO by IndustryHow to Appear in Perplexity Shopping: The New Ecommerce War in AI (2026)
Perplexity Shopping is growing faster than Google Shopping in many categories. A technical guide to get your catalog to appear and convert in AI responses.
GEO AnalyticsHow to Monitor a Product Launch in ChatGPT, Gemini and Perplexity
A practical protocol to monitor a launch in ChatGPT, Gemini and Perplexity: baseline, narrative adoption, errors, alerts and stabilization.