27 of 60: Nearly Half the Domains Leading AI Shopping Answers Haven't Cleared the Evidence Floor
Across the six published buyer-question shapes in the Machine Relations Index's Consumer Products category, 27 of the 60 top-ten citation slots belong to domains still graded 'collecting' overall — below the evidence floor that would make their rate stable. Most of the leaderboard is unclassified brand and retailer sites, not the household editorial or vendor names GEO advice assumes.
A brand team benchmarking against whoever sits at rank 1 in ChatGPT's or Gemini's shopping answers is often benchmarking against a domain the Machine Relations Index itself has not finished measuring. Across the six published buyer-question shapes in the Consumer Products category, 27 of the 60 top-ten citation slots — 45% — belong to domains whose overall confidence grade is still collecting: cited, but not yet across the evidence floor of 10 observed runs on 7 distinct dates that the Index requires before it treats a rate as stable.
This is not a claim that those domains are undeserving. It is a claim that the leaderboard a brand team screenshots today is, nearly half the time, still assembling the sample size that would make its rank mean anything next week.
The measurement
The source is the public view of the Machine Relations Index, contract machine_relations_index_public_view_v2.0, methodology mri_score_v2.0, release mri_score_v2.0+2026-09-20+47973f373a20. The machine-readable release carries every number below.
| Field | Value |
|---|---|
| Window | 2026-05-10 to 2026-09-20, 127 days observed |
| Answer runs | 16,039 |
| Citation events | 126,271 |
| Cited source domains | 22,320 |
| Engines | ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, Perplexity |
| Evidence floor per segment | 10 observed runs across 7 distinct run dates |
| Confidence grades (domain, whole-index) | A, B, C, collecting — by observation volume and consistency across the whole 127-day window, not one segment |
Consumer Products is one of the Index's 25 categories and publishes all six question shapes: three shopping shapes (best_x, top_list, x_vs_y) and three decision shapes (how_choose, is_x_worth, problem_first). Every rank below is the release's own published category_signals[].rank, read with its total beside it, never re-derived by sorting — the Index uses competition ranking, so a re-derived sort under-counts tied domains at high positions.
The top ten, shape by shape
| Shape | Segment size | Top-10 slots at collecting confidence |
|---|---|---|
| Best X | 318 domains | 2 of 10 |
| Top lists | 380 domains | 4 of 10 |
| X vs Y | 268 domains | 5 of 10 |
| How buyers choose | 427 domains | 5 of 10 |
| Is it worth it | 228 domains | 4 of 10 |
| Problem-first research | 216 domains | 7 of 10 |
Problem-first research is the worst case: 7 of its top 10 cited domains — dialavet.com, dogster.com, epa.gov, royalkennelclub.com, nopickydog.com, petco.com and canidae.com — carry no stable overall rate. Best X is the most stable shape, but even there, mealfan.com and ulta.com sit at rank 8 and 9 on collecting-grade evidence, tied at the same citation count as healthline.com, a domain the Index rates B-confidence on far more accumulated volume.
The pattern holds across shapes that look otherwise unrelated. In how_choose, koraorganics.com, cerave.com, colorescience.com, ningcos.com and brightgirl.com fill five of the top ten. In x_vs_y, wefeedraw.com, seed.com, zoopy.com, bemytravelmuse.com and nubest.com do the same. None of these domains is wrong to be cited — they are being cited, in this window, at this rate. What collecting-confidence means is that the Index has not yet observed enough distinct run dates to say the rate will hold.
It is also not the leaderboard GEO advice assumes
Most brand-visibility guidance talks about "beating Forbes" or "getting into vendor comparison lists" — editorial media and vendor-owned sources. Across the same 60 top-ten slots, the source-role breakdown is:
| Source role | Top-10 slots (of 60) |
|---|---|
| Uncategorized source (unclassified brand, retailer or niche site) | 38 |
| Editorial media | 10 |
| Platform/search (YouTube) | 6 |
| Community/social (Reddit) | 5 |
| Academic/government | 1 |
| Vendor-owned | 0 |
Sixty-three percent of the slots that actually win Consumer Products shopping and decision answers are domains the Index's nine source-role classes do not sort into any of the eight named categories — not Forbes-style editorial, not a vendor's own site, not a review platform. They are the brand and category sites themselves: whowhatwear.com, outdoorgearlab.com, chewy.com, dogfoodadvisor.com, petmd.com, walmart.com. A brand team optimizing for "get Forbes to mention us" is optimizing for 10 of 60 slots and ignoring the class that actually holds most of them.
What this means for a brand team
Two separate checks before a competitor's AI citation rank changes a strategy:
- Read the confidence grade before the rank. A rank-1 position on collecting-grade evidence is a small-sample lead, not a market position. The Index publishes both fields on the same domain record; a rank without its confidence grade is half the number.
- Look past the classified source-role classes. Most of what wins a shopping-shape answer in Consumer Products is not Forbes, not a vendor comparison page, and not a review platform — it is an unclassified brand or category site. The domain to study is the one actually sitting in the top ten, not the class of domain most GEO commentary assumes wins.
Method and sources
Every domain, rank and confidence grade above is read directly from the public Machine Relations Index release mri_score_v2.0+2026-09-20+47973f373a20: relative_signal.category_signals[] filtered to category: "consumer-products", status: "published", for rank <= 10 in each of the six shapes, cross-referenced against each domain's own mri_score_v2.overall.confidence. Nothing here is re-derived by sorting a segment on citation rate — every rank and total is the release's own published field. Correlation only: which domains the Index observed being cited, and at what confidence, over this window. It is not a claim about why, or that today's rank predicts tomorrow's.