StoreSteadyStoreSteady
Summer 2026 / Agentic Commerce Breakage Index

29 of 98 fully evaluable Top-500 Shopify stores had at least one confirmed AI-commerce breakage.

StoreSteady scanned the External Top-500 corpus, asked AI shopping agents to shop each store, and compared each answer with observable store evidence. The headline denominator is the smaller full-bundle set where every published metric had enough evidence to evaluate.

External Top-500 = Tranco VNGQN from 2026-05-21, filtered to retail Shopify-detected domains. Scans collected 2026-05-23T16:50:03.820Z to 2026-05-23T23:24:45.338Z. Agent/model IDs: claude-sonnet-4-6, gemini-2.5-flash, gpt-4o-mini, sonar-pro. Prompt hash 3af595138de3.

See whether your store has confirmed breakage or access friction

The report found confirmed breakage in 29 of 98 fully evaluable stores and access friction in 210 of 500 scanned stores. Run a store-specific scan for the same AI-shopping failure modes; the result is evidence for your store, not a projection from the benchmark.

Evidence funnel
External Top-500 stores scanned
500
Full-bundle evaluable stores
98
Stores with at least one confirmed breakage
29
Wilson 95% CI
21.5% - 39.3%
Access friction, reported separately
210 / 500
Scanned corpus
500

External Top-500 Shopify stores entered the live run.

Full-bundle evaluable
98

Stores with enough evidence across the published metric bundle.

Confirmed breakage
29

Stores with at least one calibrated and reproducible finding.

Access friction
210

Crawler, WAF, bot, or access issues. Not counted as breakage.

The denominator collapses because each detector needs observable truth: product evidence, policy text, checkout-adjacent proof, calibrated candidate labels, reproducibility, and safe non-circumventing access. Stores that block or obscure access remain visible as access friction because that is operationally important, but those stores do not inflate confirmed breakage rates.

Second finding

42.0% of the External Top-500 had access friction, reported separately from breakage.

210 of 500 scanned stores had crawler, WAF, bot, robots, or access issues that affect AI evaluability before detector-specific breakage can be confirmed. Top signals: crawl failed without classified pages (208); crawl blocked by robots without classified pages (2).

Read the confirmed set first

29 of 98 fully evaluable stores had a calibrated, reproducible breakage. Use that as the report headline, not the raw percentage alone.

Checkout remains tentative

Checkout findings are shown only with their own denominator and caveat. StoreSteady did not place orders or claim actual checkout failure.

Use this as an audit checklist

The report points to product truth, policy, offer, non-public product, checkout-path, branded-prompt, and access-friction surfaces a merchant can test on its own store.

Showing External Top-500 / All

Agent cuts are filterable data only: agent-filtered rows disclose their own denominator and CI. They are not used for launch copy or agent rankings in this report.

Published Metrics

Headline = Passed publication gates.

Tentative = Shown with its own caveat and denominator; not used in headline claims.

Not Published = Held back this quarter because calibration or evidence gates failed.

Access Friction = Reported separately; not counted as breakage.

Wrong priceHeadline

Wrong price or stock

4.4%
CI 2.5% - 7.5%

We count a product answer as wrong when price differs by more than 2% or more than $5, stock contradicts the PDP, the product does not resolve, or a material attribute conflicts with product evidence.

External Top-500
4.4%

12 broken / 273 evaluable. Alpha 92.3%. Accuracy 99.0%.

Example pattern

An agent states a product price or stock state. The live PDP or product feed shows a different value. The mismatch must reproduce before it counts.

What this means for operators
Operational risk
Product answers are only useful when the agent and the store agree on observable facts such as price, stock, product resolution, and material attributes.
What StoreSteady measured
StoreSteady compared agent product claims with current public product evidence and counted only calibrated, reproducible contradictions.
What StoreSteady did not measure
StoreSteady did not measure conversion impact, refund rate, chargeback rate, support-ticket volume, or downstream shopper behavior.
Hidden productHeadline

Surfaces a non-public product

1.8%
CI 0.8% - 4.2%

Only externally verifiable non-public signals count: 401/403, 410, noindex, robots disallow, or explicit B2B-only copy. Public absence alone is not proof.

External Top-500
1.8%

5 broken / 272 evaluable. Alpha 100.0%. Accuracy 100.0%.

Example pattern

An agent recommends a product that public evidence shows is inaccessible, noindexed, removed, or B2B-only.

What this means for operators
Operational risk
Agent answers can reference products that the public storefront marks as inaccessible, removed, noindexed, robots-blocked, or B2B-only.
What StoreSteady measured
StoreSteady counted only externally verifiable non-public signals in the public corpus, such as 401/403, 410, noindex, robots disallow, or explicit B2B-only copy.
What StoreSteady did not measure
StoreSteady did not infer Shopify draft status for external stores, bypass access controls, or measure commercial impact from exposure.
Wrong checkoutTentative

Wrong checkout path

25.0%
CI 17.7% - 34.1%

We compare payment method, guest-checkout, and shipping-country claims against public checkout, payment, help, or policy evidence.

Tentative

Checkout remains tentative unless the metric clears the headline calibration floor. StoreSteady did not place orders; checkout evidence is limited to public checkout-adjacent evidence and safe non-transactional probes.

External Top-500
25.0%

26 broken / 104 evaluable. Alpha 79.4%. Accuracy 100.0%.

Example pattern

An agent claims a payment method, guest checkout, or ship-to country. Public checkout-adjacent evidence contradicts the claim.

What this means for operators
Operational risk
Checkout-path answers are sensitive because payment methods, guest checkout, and shipping destinations must match what the store actually supports.
What StoreSteady measured
StoreSteady compared agent checkout claims with public checkout-adjacent evidence and safe non-transactional probes.
What StoreSteady did not measure
StoreSteady did not place orders, bypass checkout controls, measure checkout abandonment, or estimate conversion impact.
Not published in this report

Metrics held back by calibration gates

Wrong return or shipping rule

This metric is disclosed in the methodology but removed from headline and metric cards until the calibration gate clears.

  • Judge alpha is below 0.667, so the metric is dropped for the quarter.
  • At least one judge is below the 95% accuracy threshold.
  • Calibration accuracy is below the 95% publication threshold, so the metric is not published this quarter.
Wrong promo

This metric is disclosed in the methodology but removed from headline and metric cards until the calibration gate clears.

  • Judge alpha is below 0.667, so the metric is dropped for the quarter.
  • At least one judge is below the 95% accuracy threshold.
Competitor on your branded query

This metric is disclosed in the methodology but removed from headline and metric cards until the calibration gate clears.

  • Calibration accuracy is below the 95% publication threshold, so the metric is not published this quarter.
Metric Quality

Calibration and denominator summary

MetricTierEvaluableBreakagesRateWilson CIAlphaAccuracyCaveat
Wrong price or stockHeadline273124.4%2.5% - 7.5%92.3%99.0%Metric clears alpha, calibration accuracy, candidate mix, and false-negative gates.
Wrong return or shipping ruleNot published in this reportUnavailableUnavailableUnavailableUnavailable27.0%56.6%Judge alpha is below 0.667, so the metric is dropped for the quarter.
Wrong promoNot published in this reportUnavailableUnavailableUnavailableUnavailable18.8%99.0%Judge alpha is below 0.667, so the metric is dropped for the quarter.
Surfaces a non-public productHeadline27251.8%0.8% - 4.2%100.0%100.0%Metric clears alpha, calibration accuracy, candidate mix, and false-negative gates.
Wrong checkout pathTentative1042625.0%17.7% - 34.1%79.4%100.0%Judge alpha is below 0.8, so the metric can only publish as tentative.
Competitor on your branded queryNot published in this reportUnavailableUnavailableUnavailableUnavailable98.3%67.5%Calibration accuracy is below the 95% publication threshold, so the metric is not published this quarter.
Evidence Examples

What counted, with store identity removed

These examples are derived from normalized calibration evidence packages, not raw store crawl artifacts or raw AI responses. Store names, domains, product URLs, source quotes, exact promo codes, and merchant-identifying details are omitted before an example can render.

Wrong price or stock
store_3dac08cdb3e5
Headline
Prompt shape
Product-answer prompt asking an AI shopping agent for price, availability, and material product facts.
Agent claim
Agent claim: product exists = true.
Store evidence
Store evidence: product exists = false.
Why it counted
product exists binary claim contradicts observed truth. 3 of 3 runs reproduced.
Surfaces a non-public product
store_8bbfd3185340
Headline
Prompt shape
Product-discovery prompt asking an AI shopping agent to recommend from an anonymized candidate set.
Agent claim
Agent claim: product is public = true.
Store evidence
Store evidence: product is public = false.
Why it counted
product is public binary claim contradicts observed truth. 3 of 3 runs reproduced.
Wrong checkout path
store_b336e98a1e82
Tentative
Prompt shape
Checkout-capability prompt asking an AI shopping agent about payment, guest checkout, or ship-to support.
Agent claim
Agent claim: payment method supported:visa = true.
Store evidence
Store evidence: payment method supported:visa = false.
Why it counted
payment method supported:visa binary claim contradicts observed truth. 3 of 3 runs reproduced.
Methodology Summary

How the benchmark is counted

Store-level headline
A store counts as broken for a metric only when at least one confirmed finding survives calibration and the 2-of-3 reproducibility rule.
Measured scope
This benchmark measures answer fidelity and access friction. It does not model conversion impact, revenue loss, refunds, chargebacks, or support volume.
Metric denominators
Each metric publishes its own evaluable_store_n. Stores with missing truth stay out of that metric denominator and appear in the exclusion funnel.
Public artifacts
Public artifacts omit store names, raw URLs, and raw AI responses. The CSVs publish aggregate fields and calibration labels only.
Evidence examples
Public examples use a fail-closed safety review. If a candidate package contains unsafe identifiers, the report shows why the example is unavailable instead of weakening the redaction rules.
Store Check

The useful next step is not a market claim. It is checking whether your own store creates these same AI-shopping failure modes.

The report found confirmed breakage in 29 of 98 fully evaluable stores and access friction in 210 of 500 scanned stores. Run a store-specific scan for the same AI-shopping failure modes; the result is evidence for your store, not a projection from the benchmark.

See whether your store has confirmed breakage or access friction