29 of 98 fully evaluable Top-500 Shopify stores had at least one confirmed AI-commerce breakage.
StoreSteady scanned the External Top-500 corpus, asked AI shopping agents to shop each store, and compared each answer with observable store evidence. The headline denominator is the smaller full-bundle set where every published metric had enough evidence to evaluate.
External Top-500 = Tranco VNGQN from 2026-05-21, filtered to retail Shopify-detected domains. Scans collected 2026-05-23T16:50:03.820Z to 2026-05-23T23:24:45.338Z. Agent/model IDs: claude-sonnet-4-6, gemini-2.5-flash, gpt-4o-mini, sonar-pro. Prompt hash 3af595138de3.
See whether your store has confirmed breakage or access frictionThe report found confirmed breakage in 29 of 98 fully evaluable stores and access friction in 210 of 500 scanned stores. Run a store-specific scan for the same AI-shopping failure modes; the result is evidence for your store, not a projection from the benchmark.
- External Top-500 stores scanned
- 500
- Full-bundle evaluable stores
- 98
- Stores with at least one confirmed breakage
- 29
- Wilson 95% CI
- 21.5% - 39.3%
- Access friction, reported separately
- 210 / 500
External Top-500 Shopify stores entered the live run.
Stores with enough evidence across the published metric bundle.
Stores with at least one calibrated and reproducible finding.
Crawler, WAF, bot, or access issues. Not counted as breakage.
The denominator collapses because each detector needs observable truth: product evidence, policy text, checkout-adjacent proof, calibrated candidate labels, reproducibility, and safe non-circumventing access. Stores that block or obscure access remain visible as access friction because that is operationally important, but those stores do not inflate confirmed breakage rates.
42.0% of the External Top-500 had access friction, reported separately from breakage.
210 of 500 scanned stores had crawler, WAF, bot, robots, or access issues that affect AI evaluability before detector-specific breakage can be confirmed. Top signals: crawl failed without classified pages (208); crawl blocked by robots without classified pages (2).
29 of 98 fully evaluable stores had a calibrated, reproducible breakage. Use that as the report headline, not the raw percentage alone.
Checkout findings are shown only with their own denominator and caveat. StoreSteady did not place orders or claim actual checkout failure.
The report points to product truth, policy, offer, non-public product, checkout-path, branded-prompt, and access-friction surfaces a merchant can test on its own store.
Showing External Top-500 / All
Agent cuts are filterable data only: agent-filtered rows disclose their own denominator and CI. They are not used for launch copy or agent rankings in this report.
Published Metrics
Headline = Passed publication gates.
Tentative = Shown with its own caveat and denominator; not used in headline claims.
Not Published = Held back this quarter because calibration or evidence gates failed.
Access Friction = Reported separately; not counted as breakage.
Wrong price or stock
We count a product answer as wrong when price differs by more than 2% or more than $5, stock contradicts the PDP, the product does not resolve, or a material attribute conflicts with product evidence.
12 broken / 273 evaluable. Alpha 92.3%. Accuracy 99.0%.
An agent states a product price or stock state. The live PDP or product feed shows a different value. The mismatch must reproduce before it counts.
- Operational risk
- Product answers are only useful when the agent and the store agree on observable facts such as price, stock, product resolution, and material attributes.
- What StoreSteady measured
- StoreSteady compared agent product claims with current public product evidence and counted only calibrated, reproducible contradictions.
- What StoreSteady did not measure
- StoreSteady did not measure conversion impact, refund rate, chargeback rate, support-ticket volume, or downstream shopper behavior.
Wrong checkout path
We compare payment method, guest-checkout, and shipping-country claims against public checkout, payment, help, or policy evidence.
Checkout remains tentative unless the metric clears the headline calibration floor. StoreSteady did not place orders; checkout evidence is limited to public checkout-adjacent evidence and safe non-transactional probes.
26 broken / 104 evaluable. Alpha 79.4%. Accuracy 100.0%.
An agent claims a payment method, guest checkout, or ship-to country. Public checkout-adjacent evidence contradicts the claim.
- Operational risk
- Checkout-path answers are sensitive because payment methods, guest checkout, and shipping destinations must match what the store actually supports.
- What StoreSteady measured
- StoreSteady compared agent checkout claims with public checkout-adjacent evidence and safe non-transactional probes.
- What StoreSteady did not measure
- StoreSteady did not place orders, bypass checkout controls, measure checkout abandonment, or estimate conversion impact.
Metrics held back by calibration gates
This metric is disclosed in the methodology but removed from headline and metric cards until the calibration gate clears.
- Judge alpha is below 0.667, so the metric is dropped for the quarter.
- At least one judge is below the 95% accuracy threshold.
- Calibration accuracy is below the 95% publication threshold, so the metric is not published this quarter.
This metric is disclosed in the methodology but removed from headline and metric cards until the calibration gate clears.
- Judge alpha is below 0.667, so the metric is dropped for the quarter.
- At least one judge is below the 95% accuracy threshold.
This metric is disclosed in the methodology but removed from headline and metric cards until the calibration gate clears.
- Calibration accuracy is below the 95% publication threshold, so the metric is not published this quarter.
Calibration and denominator summary
| Metric | Tier | Evaluable | Breakages | Rate | Wilson CI | Alpha | Accuracy | Caveat |
|---|---|---|---|---|---|---|---|---|
| Wrong price or stock | Headline | 273 | 12 | 4.4% | 2.5% - 7.5% | 92.3% | 99.0% | Metric clears alpha, calibration accuracy, candidate mix, and false-negative gates. |
| Wrong return or shipping rule | Not published in this report | Unavailable | Unavailable | Unavailable | Unavailable | 27.0% | 56.6% | Judge alpha is below 0.667, so the metric is dropped for the quarter. |
| Wrong promo | Not published in this report | Unavailable | Unavailable | Unavailable | Unavailable | 18.8% | 99.0% | Judge alpha is below 0.667, so the metric is dropped for the quarter. |
| Surfaces a non-public product | Headline | 272 | 5 | 1.8% | 0.8% - 4.2% | 100.0% | 100.0% | Metric clears alpha, calibration accuracy, candidate mix, and false-negative gates. |
| Wrong checkout path | Tentative | 104 | 26 | 25.0% | 17.7% - 34.1% | 79.4% | 100.0% | Judge alpha is below 0.8, so the metric can only publish as tentative. |
| Competitor on your branded query | Not published in this report | Unavailable | Unavailable | Unavailable | Unavailable | 98.3% | 67.5% | Calibration accuracy is below the 95% publication threshold, so the metric is not published this quarter. |
What counted, with store identity removed
These examples are derived from normalized calibration evidence packages, not raw store crawl artifacts or raw AI responses. Store names, domains, product URLs, source quotes, exact promo codes, and merchant-identifying details are omitted before an example can render.
- Prompt shape
- Product-answer prompt asking an AI shopping agent for price, availability, and material product facts.
- Agent claim
- Agent claim: product exists = true.
- Store evidence
- Store evidence: product exists = false.
- Why it counted
- product exists binary claim contradicts observed truth. 3 of 3 runs reproduced.
- Prompt shape
- Product-discovery prompt asking an AI shopping agent to recommend from an anonymized candidate set.
- Agent claim
- Agent claim: product is public = true.
- Store evidence
- Store evidence: product is public = false.
- Why it counted
- product is public binary claim contradicts observed truth. 3 of 3 runs reproduced.
- Prompt shape
- Checkout-capability prompt asking an AI shopping agent about payment, guest checkout, or ship-to support.
- Agent claim
- Agent claim: payment method supported:visa = true.
- Store evidence
- Store evidence: payment method supported:visa = false.
- Why it counted
- payment method supported:visa binary claim contradicts observed truth. 3 of 3 runs reproduced.
How the benchmark is counted
- Store-level headline
- A store counts as broken for a metric only when at least one confirmed finding survives calibration and the 2-of-3 reproducibility rule.
- Measured scope
- This benchmark measures answer fidelity and access friction. It does not model conversion impact, revenue loss, refunds, chargebacks, or support volume.
- Metric denominators
- Each metric publishes its own evaluable_store_n. Stores with missing truth stay out of that metric denominator and appear in the exclusion funnel.
- Public artifacts
- Public artifacts omit store names, raw URLs, and raw AI responses. The CSVs publish aggregate fields and calibration labels only.
- Evidence examples
- Public examples use a fail-closed safety review. If a candidate package contains unsafe identifiers, the report shows why the example is unavailable instead of weakening the redaction rules.
The useful next step is not a market claim. It is checking whether your own store creates these same AI-shopping failure modes.
The report found confirmed breakage in 29 of 98 fully evaluable stores and access friction in 210 of 500 scanned stores. Run a store-specific scan for the same AI-shopping failure modes; the result is evidence for your store, not a projection from the benchmark.
See whether your store has confirmed breakage or access friction