Product Availability Contradiction Protocol for AI Shopping Answers
A preregistered matrix for testing whether AI shopping answers separate product existence from buyer-region and observation-time availability.
A product-availability contradiction test asks a narrow question: when an AI shopping answer recommends a product, did it distinguish the product's existence from whether that product was actually available for the buyer's region at the observation time? The test should lock product identity, region, timestamp, first-party evidence, answer claim, cited support, and an explicit unknown label before anyone reads the result.
This is an unexecuted Para Labs protocol. It is not a finding about ChatGPT, Google, a retailer, a marketplace, or any current shopping-answer provider. It is a preregistration template and blank scoring rubric for teams that need to test availability claims without inventing certainty after the answer appears.
Product existence is not the same as buyer-region availability
The failure mode is not merely that an AI system names the wrong product. It is that the product may exist, match the requested identity, and still be unavailable for the buyer who asked the question. OpenAI's product-discovery announcement frames shopping in ChatGPT around comparing options with details such as price, reviews, and features, and says Agentic Commerce Protocol support is meant to bring more complete and up-to-date product information into ChatGPT (OpenAI). OpenAI's merchant page separately tells merchants that product feeds can keep pricing, availability, and updates current, while noting that shopping is live for ChatGPT users in the U.S. and that merchant data integration is expanding over time (ChatGPT merchants).
That makes availability a testable claim, not a vibes-based shopping impression. A recommendation can be wrong in at least four different ways:
| Error class | What the answer gets wrong | Why ordinary product-name checking misses it |
|---|---|---|
| Nonexistent product | The named product cannot be found in first-party or reliable retail evidence | Identity validation catches this |
| Wrong variant | The family exists, but the answer points to the wrong size, color, generation, bundle, or SKU | Product-name matching often collapses variants |
| Region mismatch | The product exists somewhere, but not in the buyer's shipping, pickup, or store region | A global product page may look valid while local fulfillment fails |
| Timestamp mismatch | The product was available before or after the observation window, but not when the answer made the claim | Cached pages and stale snippets can preserve old availability |
The September 12 Para Labs protocol on same-product-name variant identity already covers the second row. This protocol deliberately holds identity fixed and tests the availability rows: region and time.
Preregister the availability matrix before running prompts
The preregistration should define a product as eligible only when its identity can be fixed separately from the availability treatment. Do not choose a product because it will create an interesting failure. Choose it because the first-party or merchant evidence can identify the product and its availability state at a specific region and timestamp.
Use this blank matrix before collecting any AI answer:
| Field | Locked value to fill before testing | Notes for annotators |
|---|---|---|
| Product identity | Brand, product family, exact model, SKU/GTIN if available, variant attributes, product URL | Identity is fixed before availability is judged |
| Buyer region | Country, state/province, postal code or store radius, and fulfillment mode | Availability must be scored for the buyer context, not globally |
| Observation timestamp | ISO timestamp, timezone, and page-capture time for every evidence source | The answer is judged against the evidence available at observation time |
| First-party evidence | Manufacturer page, merchant page, product feed field, pickup/delivery page, or archived capture | Prefer direct product and merchant evidence over roundups |
| Availability state | In stock, out of stock, preorder, backorder, limited availability, unknown, or not supported by evidence | The label belongs to the buyer-region/timestamp pair |
| AI answer claim | Exact answer text, product name, availability wording, price if stated, and caveats | Record the claim before judging correctness |
| Cited support | URLs, citation snippets, source dates, and whether each source supports availability at the required granularity | A product review does not prove current stock |
| Unknown rule | Conditions under which the answer receives an explicit unknown instead of correct/incorrect | Unknown is a valid label, not a failed annotation |
Google's Merchant Center documentation treats availability as a required product-data attribute and says values should match the website; it defines states such as in stock, out of stock, preorder, and backorder (Google Merchant Center). Schema.org's availability property is the web markup vocabulary for expressing states such as in stock and out of stock inside an offer (Schema.org). Those standards do not tell us whether an AI answer used the data correctly. They tell us what evidence an annotator can look for before scoring the answer.
The contradiction cells should vary region and time, not product identity
A clean availability test keeps the product constant and changes only the availability context. If the product identity changes between cells, the experiment has become another variant-identity test. If the source set changes between cells, it has become a source-ablation test. Availability contradiction needs tighter controls.
Use these cells:
| Cell | Product identity | Region and time condition | Expected annotation question |
|---|---|---|---|
| Baseline available | Same fixed product | Region A, timestamp T, first-party evidence says available | Does the answer state availability with support? |
| Baseline unavailable | Same fixed product | Region B or timestamp T, evidence says unavailable | Does the answer avoid recommending it as available? |
| Regional split | Same fixed product | Region A available, Region B unavailable at the same timestamp | Does the answer preserve the buyer's region? |
| Temporal split | Same fixed product | Available at T1, unavailable at T2, or reverse | Does the answer preserve the observation timestamp? |
| Ambiguous evidence | Same fixed product | Product page exists, but local stock or shipping status is absent | Does the answer label availability as unknown? |
| Citation mismatch | Same fixed product | Cited source supports identity or features but not availability | Does the answer avoid treating weak support as stock proof? |
The final cell is often the most important. OpenAI's shopping research article says the feature can look across the internet for up-to-date information such as price and availability, while also warning that shopping research may make mistakes about product details including price and availability and encouraging users to visit merchant sites for the most accurate details (OpenAI shopping research). A protocol should therefore score the support relationship, not merely whether the answer sounds current.
Blank scoring rubric for availability claims
Score product identity, availability, and source support as separate dimensions. An answer can identify the correct product but overstate stock. It can cite a merchant page but ignore region. It can be directionally useful while still failing the availability claim.
| Dimension | Label | Definition | Evidence required |
|---|---|---|---|
| Product identity | Correct product | Brand, model, and variant match the preregistered product | First-party or merchant page confirms the exact product |
| Product identity | Wrong variant | Product family matches but material variant differs | SKU, model, generation, size, color, bundle, or specs conflict |
| Product identity | Unresolved identity | Evidence cannot confirm the exact product | Missing SKU/model or conflicting pages |
| Buyer-region availability | Correct available | Answer says available and evidence supports availability for the locked region and timestamp | Merchant, feed, or page capture supports the buyer context |
| Buyer-region availability | Correct unavailable | Answer withholds or caveats recommendation because evidence says unavailable | Evidence supports unavailability or non-fulfillment |
| Buyer-region availability | Contradiction | Answer says available where the evidence says unavailable, or the reverse | Same product, same region, same timestamp |
| Buyer-region availability | Unsupported availability | Answer states availability when source only supports identity, features, or historical listing | Citation does not entail stock/fulfillment |
| Buyer-region availability | Unknown handled | Answer labels availability as unknown or tells the user to verify at merchant site when evidence is missing | Absence of granular evidence is preserved |
| Source support | Supported | Cited source entails the exact availability claim | Source covers product, region, fulfillment mode, and timestamp |
| Source support | Partially supported | Source supports product identity or a weaker availability statement | Unsupported part is named |
| Source support | Unsupported | Source does not support the availability claim | Review, listicle, stale page, or nonregional evidence is used as proof |
Do not collapse these into one pass/fail score. The useful result may be: correct product, unsupported availability, citation mismatch. That is more actionable than simply marking the answer wrong.
Reporting should separate protocol output from market demand
The demand signal for this specific Para Labs asset is weak and should stay weak. The adjacent Search Console evidence behind this protocol was only four blog query-page rows, nine impressions, and zero clicks. Three qualified queries tied to ChatGPT product discovery or product feeds showed one impression each at the existing Para Labs product-discovery article; six impressions were from a site:vercel.app query, which is not qualified shopping demand. The crawl evidence showed assistant and crawler visits, not buyers.
The better reason to publish the protocol is methodological fit. Shopping-answer claims combine identity, source, region, and time in a way that ordinary AI visibility dashboards can blur. A falsifiable blank rubric gives teams a way to investigate that blur without pretending the crawl logs prove commercial demand.
The current Machine Relations Index is the right data authority for the broader citation-occurrence context, not for retail stock truth. Its September 14, 2026 release reports enterprise-software evidence from 15,396 observed answer runs across a May 10-September 14 window and includes an example in which erpresearch.com appeared in 99 of 1,237 enterprise-software runs. That shows market databases can enter answer evidence. It does not prove a cited product page establishes real-time buyer availability.
Machine Relations implication: availability is a claim-support problem
Machine Relations treats AI visibility as a source-condition problem: machines answer from the evidence they can retrieve, trust, and represent. Product availability raises the same issue at a more volatile layer. The presence of a product citation is not enough. The source must support the specific claim the answer makes for the buyer's context.
For commerce teams, the operating sequence is practical: maintain product identity data, expose availability in machine-readable fields, capture first-party evidence at the time of observation, run fixed-identity contradiction tests, and report unknowns when the source cannot support a region/time claim. Google's structured-data guidance says supported merchant markup includes price, price currency, availability, and condition for automatic item updates (Google Merchant Center structured data). That is infrastructure for measurement; it is not a guarantee that every answer engine will preserve the claim correctly.
AuthorityTech can use this kind of protocol inside AI visibility work because it keeps the evidence burden in the open. A brand does not need a dramatic hallucination story. It needs to know which claims are supported, which are contradicted, and which should be labeled unknown before anyone turns an AI answer into a forecast.
Teams that want a faster outside map before designing tests can run an AI visibility audit, then reserve availability contradiction tests for the products and regions where an incorrect recommendation would materially affect the buyer experience.
FAQ
What is a product-availability contradiction test?
A product-availability contradiction test checks whether an AI shopping answer preserves the difference between a product that exists and a product that is available for a specific buyer region at a specific observation time. It fixes product identity first, then scores availability and cited support separately.
Why is the unknown label necessary?
Unknown prevents an unsupported source from being forced into correct or incorrect. If a cited page proves the product exists but does not prove local stock, shipping eligibility, pickup status, or the observation timestamp, the honest label is unknown or unsupported availability.
Does product-feed availability prove an AI answer is correct?
No. A feed or structured-data field can be strong evidence for availability when it matches the website and buyer context, but the AI answer still has to preserve the correct product, region, timestamp, and support relationship. The protocol tests that preservation.
How is this different from source ablation?
Source ablation varies the evidence source to test answer sensitivity. This protocol holds both product identity and source plan as stable as possible, then varies buyer-region and timestamp availability. The question is not whether a source caused the answer; it is whether the answer supported the availability claim it made.