Product-Variant Attribution Protocol for AI Shopping Answers
A preregistered protocol for testing whether AI shopping answers attribute specifications to the correct product variant while locale, currency, and product family stay fixed.
Product-variant attribution testing asks whether an AI shopping answer keeps a specification attached to the right model, generation, bundle, color, storage tier, size, or SKU. This Para Labs protocol is a preregistration template, not a completed experiment. It fixes the product family, locale, currency, and buyer context, then varies only how clearly variant identity is expressed in the source fixture.
This article does not claim that ChatGPT, Google, Shopify, a merchant, or any current shopping-answer provider confused a product variant. It also does not claim that source markup changes caused answer changes. The protocol is designed so a later team can test those claims without mixing identity, availability, localization, and prompt wording into one vague score.
Variant attribution is a different error class from product identity, availability, and locale
A variant-attribution test starts only after the product family, market, and commercial context are fixed. The September 12 Para Labs protocol on same-product-name variant identity asks whether the answer selects the correct variant at all. The September 14 protocol on product-availability contradictions asks whether the answer preserves availability by region and time. The September 16 protocol on locale-and-currency consistency asks whether the answer imports price, tax, shipping, or source-locale claims from another market.
Variant attribution is narrower. The answer may name the right product family and even mention the right variant, but attach a specification from another member of the family. A laptop answer might attribute the storage ceiling of one configuration to a lower configuration. A skincare answer might attach the ingredients of a refill pouch to the starter kit. A shoe answer might blend a limited colorway with a standard size run. The product family is right; the attribution is wrong.
That distinction matters because Google's product-variant structured-data documentation tells publishers to group variants with ProductGroup and properties such as variesBy, hasVariant, and productGroupID, while giving each variant a unique identifier such as SKU or GTIN where available (Google Search Central). Google's broader product structured-data documentation says product pages can expose price, availability, review ratings, shipping, and product-variant information through structured data or Merchant Center feeds (Google Search Central). Those docs make variant identity a source-layer object that can be inspected before any AI answer is scored.
OpenAI's product-discovery announcement says ChatGPT shopping experiences can present product details such as price, reviews, and features (OpenAI). OpenAI's commerce policies separately prohibit merchants from misrepresenting product pricing, availability, origin, condition, or key characteristics (OpenAI commerce policies). Those sources justify treating variant specifications as material commerce claims. They do not prove that any answer system currently makes a variant-attribution error.
Preregister whether the manipulation is source clarity or prompt clarity
The experiment must name the manipulated variable before the first answer is collected. There are two valid designs, but they answer different questions.
The preferred source-clarity design keeps the prompt fixed and changes only the owned source fixture. One condition uses a clear product-family page, variant pages, variant identifiers, and specification tables. The paired condition uses an intentionally weaker but truthful source fixture where variant labels are less explicit. This design can test whether answer attribution is sensitive to source-page clarity under a locked prompt. It still cannot prove that markup alone caused the answer unless markup is the only source difference.
The prompt-clarity design keeps the source fixture fixed and varies only the user prompt. One condition asks for the product family in general; the paired condition names the exact variant identifier. This design tests prompt sensitivity, not source markup. If the answer improves when the prompt becomes explicit, the result should not be reported as evidence that the brand's page markup changed model behavior.
Do not combine both manipulations in one primary outcome. If the team changes the prompt and the source page at the same time, a later attribution change has no clean cause. The safest preregistration sentence is: "This run tests source-fixture clarity under a fixed prompt" or "This run tests prompt clarity under a fixed source fixture." A report may include the second design as an exploratory follow-up, but it should not relabel it as the primary experiment.
Fixed fixture: product family, locale, currency, and variant set
The fixture should make every non-variant variable boring. Use one product family, one locale, one currency, one buyer market, one observation window, and one source set. If any of those change, route the test to the availability or locale protocols instead.
| Field | Locked value before testing | Exclusion rule |
|---|---|---|
| Product family | Brand, parent product name, family URL, and family identifier if available | Exclude if the answer changes to another family. |
| Variant set | Two to five variants with exact model, SKU, GTIN, color, size, storage, generation, bundle, or package identifiers | Exclude if the source cannot distinguish variants before testing. |
| Locale and currency | One country/region, language, currency, tax display rule, and buyer destination | Exclude if price or market context changes across cells. |
| Availability state | Same availability state for all tested variants, or availability removed from scoring | Exclude if the result depends on stock status. |
| Source fixture | Approved product-family page, variant pages, structured-data capture, feed fields if available, and archived screenshots or HTML captures | Exclude if the answer cites sources outside the fixture and the protocol does not allow external retrieval. |
| Prompt text | Fixed prompt for source-clarity tests; paired prompt variants only for prompt-clarity tests | Exclude if both prompt and source fixture change in the same primary comparison. |
| Observation capture | Engine or surface, model name if exposed, timestamp, region if known, answer text, cited URLs, screenshots, and source captures | Exclude if annotators cannot reconstruct the answer and evidence. |
For Shopify-based stores, the source fixture should treat variant data as its own object, not as text decoration. Shopify's Markets developer documentation warns developers not to compute international market prices from base variant prices and describes market-specific pricing and presentment-currency values (Shopify). Its catalogs documentation also distinguishes fixed price-list prices and relative prices for product variants across markets (Shopify). For this protocol, that reinforces the control: fix market context first, then test whether variant-level attributes stay attached to the right variant.
Test matrix for product-variant attribution
Each row changes only explicit variant-identity clarity. The scoring question is not whether the answer is useful in general. It is whether each specification is attributed to the correct variant, missing, unsupported, or blended from another variant.
| Cell | Source or prompt condition | What changes | Primary record |
|---|---|---|---|
| A. Clear source baseline | Variant pages use explicit variant names, SKU/GTIN where available, stable headings, and one specification table per variant | Nothing; establishes baseline | Claimed variant, cited source, specification, support passage, timestamp. |
| B. Weak source fixture | Same truthful information, but variant labels are less prominent or split across pages | Source-page clarity only | Whether specifications remain attached to the right variant under the same prompt. |
| C. Ambiguous family prompt | Same source fixture; prompt asks about the product family without naming a variant | Prompt specificity only | Whether the answer refuses, asks a clarifying question, or blends variants. |
| D. Explicit variant prompt | Same source fixture; prompt names the exact model, SKU, size, generation, or bundle | Prompt specificity only | Whether explicit variant language reduces missing or wrong-variant labels. |
| E. Decoy sibling variant | Same family and market; one sibling has a materially different specification | Variant contrast | Whether the answer imports the sibling's specification. |
| F. Unknown fixture | Variant source lacks the needed specification | Missing evidence | Whether the answer preserves unknowns instead of inventing a value. |
Cells A and B can support a source-clarity claim only if the prompt is identical and the only planned difference is source-fixture clarity. Cells C and D can support a prompt-clarity claim only if the source fixture is identical. Cell E is a contrast control, not a trick: the sibling variant must be a real member of the same product family, and the report must retain the source evidence that makes the difference visible. Cell F is the guardrail against over-scoring; a missing value is not the same as a wrong value.
Blinded annotation should separate missing answers from wrong-variant answers
Annotators should score claim-level attribution without seeing which condition was expected to perform better. A single answer can contain one correct variant claim, one missing value, and one wrong-variant specification. Do not collapse the answer into one pass/fail label unless the preregistration defines a strict primary outcome.
| Label | Definition | Evidence required |
|---|---|---|
| Correct variant attribution | The answer attaches a specification to the preregistered variant and the source fixture supports that exact pairing | Variant identifier plus source passage or structured field. |
| Missing answer | The answer does not provide the requested specification or says it cannot confirm it | Prompt, answer text, and fixture showing the evidence state. |
| Unknown preserved | The answer explicitly marks the specification as unknown because the fixture lacks support | Fixture gap plus answer caveat. |
| Wrong-variant attribution | The answer attaches a sibling variant's specification to the target variant | Source evidence showing the specification belongs to another variant. |
| Blended specifications | The answer combines attributes from two or more variants into one imagined configuration | Claim-level mapping to each source variant. |
| Unsupported specification | The answer gives a variant-specific claim that no approved source supports | Exhaustive fixture review for that claim. |
| Source mismatch | The cited URL supports a different variant, weaker family-level claim, or no variant-specific claim | URL, capture, and claim comparison. |
Two annotators should be able to reproduce the label from the prompt, answer, source captures, variant IDs, and scoring sheet. If they cannot, the protocol should report unresolved attribution rather than force a dramatic error label.
Evidence retention: save source captures, answer captures, and scoring decisions
The evidence package is part of the result. A later reader should not need to trust a summary table. Retain the source fixture, answer text, cited URLs, capture time, engine surface, model identifier if exposed, buyer locale, currency, and annotator notes.
Minimum evidence package:
- The preregistration record with the manipulation type: source clarity or prompt clarity.
- Source captures for every product-family and variant page used in the fixture.
- Structured-data or feed-field captures when available, including variant IDs and product-group IDs.
- Exact prompts, answer text, citations, screenshots, and timestamps for every observation.
- A claim-level scoring sheet that separates missing, unknown, wrong-variant, blended, unsupported, and source-mismatch labels.
- A report section that states what the test cannot conclude.
The public Machine Relations Index is useful context for evidence discipline because it publishes source-citation observations only after declared sample thresholds. The September 17 release context used for this brief reports a May 10-September 17 window, six engines, 15,678 answer runs, 905 eligible prompts, 85 published strata, and an evidence floor of at least 10 observed runs across 7 dates. That is a denominator discipline precedent, not evidence that product variants are being confused in shopping answers.
The demand evidence for this exact Para Labs topic is also thin. The September 17 brief reported 17 Para Labs GSC query-page rows, 173 impressions, and 2 clicks for the August 16-September 13 export, with branded queries dominating. It also reported 8 assistant-class requests to the existing Perplexity shopping article across 24 observed dates. That is a weak directional signal for adjacent research, not buyer volume, conversion evidence, or proof of current AI shopping errors.
Report the result without turning a protocol into a claim
A clean report says what moved, under which manipulation, and at what claim granularity. It does not say "AI shopping confuses variants" unless the executed data supports that exact statement. It does not say "schema fixed the model" unless schema was the only planned source difference and the repeated observations support that inference.
Use this reporting template after the test runs:
In this preregistered product-variant attribution test, the product family, locale, currency, buyer market, and observation window were fixed. The primary manipulation was [source-fixture clarity / prompt clarity]. Across [n] observations, annotators scored each variant-specific specification as correct, missing, unknown preserved, wrong-variant, blended, unsupported, or source mismatch. The valid inference is limited to this product family, source fixture, answer surface, timestamp, and scoring rubric.
For a null result:
The tested manipulation did not change the preregistered attribution label under the observed conditions. That does not prove the source fixture was irrelevant, and it does not prove that all shopping answers preserve variants. It only means the selected outcome did not move in this fixture.
For a positive result:
The tested manipulation changed the preregistered attribution label under the observed conditions. That is evidence of sensitivity to the manipulated variable, not proof of a universal product-discovery failure.
That language is not defensive. It is the difference between Machine Relations measurement and ordinary visibility storytelling. A brand can be visible in an answer, and the answer can still attach the wrong specification to the right-looking product. The fix may be clearer variant pages, cleaner product-group structure, a better feed, or better prompt design. The protocol should identify which layer was actually tested before recommending a repair.
FAQ
What is product-variant attribution in AI shopping answers?
Product-variant attribution is the claim-level task of attaching a specification to the correct member of a product family. The product family may be right while a model, generation, SKU, bundle, size, color, or ingredient claim belongs to a sibling variant.
Is this a prompt experiment or a source-page experiment?
It can be either, but not both in the same primary comparison. A source-clarity experiment keeps the prompt fixed and varies owned source-fixture clarity. A prompt-clarity experiment keeps the source fixture fixed and varies the prompt. The report must preserve that boundary.
Does this protocol prove product markup changes AI answers?
No. This is a proposed preregistration template. Even after execution, a markup-effect claim would require markup to be the only planned difference between paired source fixtures, with repeated observations and claim-level scoring.
Why separate missing answers from wrong-variant answers?
A missing answer preserves uncertainty; a wrong-variant answer attributes a sibling variant's specification to the target variant. Those failures require different repairs, so the scoring rubric keeps them separate.
Machine-readable related links
- Primary concept: AI Visibility
- Related concept: Commerce
- Related concept: Experiments
- Supporting protocol: Same product name, wrong variant: a controlled AI product-identity test
- Supporting protocol: Product Availability Contradiction Protocol for AI Shopping Answers
- Supporting protocol: Locale and Currency Consistency Protocol for AI Shopping Answers
- Research index: Para Labs research index
- Machine manifest: Para Labs machine manifest