Agentic Commerce Breakage Index Methodology
- Methodology version
- 2026-summer.public-v1
- Prompt manifest hash
- 3af595138de35b2a5655e8a2234ebdeddc915a46be62b517e73048e89e382d40
- Canonical path
- /resources/research/agentic-commerce-breakage-index/methodology
Why This Report Exists
The Breakage Index measures whether AI shopping answers match real store evidence. It is a fidelity benchmark, not an AI visibility ranking.
A store counts as broken for a metric only when at least one confirmed finding survives reproducibility and calibration gates.
Frozen Prompt Manifest
Public release: Summer 2026. Frozen prompt content hash: 3af595138de35b2a5655e8a2234ebdeddc915a46be62b517e73048e89e382d40.
The public prompt table preserves exact prompt text, model IDs, parser schemas, and audit fields; internal protocol IDs remain available through the run provenance.
System prompt: You are evaluating one public ecommerce store for StoreSteady. Answer only from the store evidence and current AI shopping answer context supplied in the prompt. If evidence is missing, say it is missing. Do not infer private, draft, or admin-only state from public absence.
Scan window: 2026-05-23T16:50:03.820Z to 2026-05-23T23:24:45.338Z.
Response audit fields: prompt_version, system_prompt_version, model_id, provider, agent_surface, timestamp, raw_response_hash, parsed_response, parser_schema_version.
| prompt_family | metric | agent | provider | model_id | system_prompt_id | prompt_version | system_prompt_version | parser_schema | required_store_fields | parser_output_fields | prompt_text |
|---|---|---|---|---|---|---|---|---|---|---|---|
| product_truth_mismatch.chatgpt.v1 | product_truth_mismatch | chatgpt | openai | gpt-4o-mini | breakage-index-shopper-system | 1 | 1 | breakage-index.product_truth_mismatch.parser.v1@1 | store_domain, store_name, primary_category, product_titles, product_urls | product_name, claimed_price, claimed_availability, claimed_attributes, source_quote | For {{store_name}} ({{store_domain}}), answer as a shopper comparing these products: {{product_titles}}. Name current price, availability, and any material product attributes you rely on. Cite uncertainty instead of guessing. |
| product_truth_mismatch.perplexity.v1 | product_truth_mismatch | perplexity | perplexity | sonar-pro | breakage-index-shopper-system | 1 | 1 | breakage-index.product_truth_mismatch.parser.v1@1 | store_domain, store_name, primary_category, product_titles, product_urls | product_name, claimed_price, claimed_availability, claimed_attributes, source_quote | For {{store_name}} ({{store_domain}}), answer as a shopper comparing these products: {{product_titles}}. Name current price, availability, and any material product attributes you rely on. Cite uncertainty instead of guessing. |
| product_truth_mismatch.claude.v1 | product_truth_mismatch | claude | anthropic | claude-sonnet-4-6 | breakage-index-shopper-system | 1 | 1 | breakage-index.product_truth_mismatch.parser.v1@1 | store_domain, store_name, primary_category, product_titles, product_urls | product_name, claimed_price, claimed_availability, claimed_attributes, source_quote | For {{store_name}} ({{store_domain}}), answer as a shopper comparing these products: {{product_titles}}. Name current price, availability, and any material product attributes you rely on. Cite uncertainty instead of guessing. |
| product_truth_mismatch.gemini.v1 | product_truth_mismatch | gemini | gemini-2.5-flash | breakage-index-shopper-system | 1 | 1 | breakage-index.product_truth_mismatch.parser.v1@1 | store_domain, store_name, primary_category, product_titles, product_urls | product_name, claimed_price, claimed_availability, claimed_attributes, source_quote | For {{store_name}} ({{store_domain}}), answer as a shopper comparing these products: {{product_titles}}. Name current price, availability, and any material product attributes you rely on. Cite uncertainty instead of guessing. | |
| policy_misstatement.chatgpt.v1 | policy_misstatement | chatgpt | openai | gpt-4o-mini | breakage-index-shopper-system | 1 | 1 | breakage-index.policy_misstatement.parser.v1@1 | store_domain, store_name, policy_urls | return_window_days, return_cost, free_shipping_threshold, ships_internationally, source_quote | For {{store_name}} ({{store_domain}}), state the return window, return cost, free-shipping threshold, and whether the store ships internationally. Use only policy evidence from these URLs: {{policy_urls}}. |
| policy_misstatement.perplexity.v1 | policy_misstatement | perplexity | perplexity | sonar-pro | breakage-index-shopper-system | 1 | 1 | breakage-index.policy_misstatement.parser.v1@1 | store_domain, store_name, policy_urls | return_window_days, return_cost, free_shipping_threshold, ships_internationally, source_quote | For {{store_name}} ({{store_domain}}), state the return window, return cost, free-shipping threshold, and whether the store ships internationally. Use only policy evidence from these URLs: {{policy_urls}}. |
| policy_misstatement.claude.v1 | policy_misstatement | claude | anthropic | claude-sonnet-4-6 | breakage-index-shopper-system | 1 | 1 | breakage-index.policy_misstatement.parser.v1@1 | store_domain, store_name, policy_urls | return_window_days, return_cost, free_shipping_threshold, ships_internationally, source_quote | For {{store_name}} ({{store_domain}}), state the return window, return cost, free-shipping threshold, and whether the store ships internationally. Use only policy evidence from these URLs: {{policy_urls}}. |
| policy_misstatement.gemini.v1 | policy_misstatement | gemini | gemini-2.5-flash | breakage-index-shopper-system | 1 | 1 | breakage-index.policy_misstatement.parser.v1@1 | store_domain, store_name, policy_urls | return_window_days, return_cost, free_shipping_threshold, ships_internationally, source_quote | For {{store_name}} ({{store_domain}}), state the return window, return cost, free-shipping threshold, and whether the store ships internationally. Use only policy evidence from these URLs: {{policy_urls}}. | |
| offer_ambiguity.chatgpt.v1 | offer_ambiguity | chatgpt | openai | gpt-4o-mini | breakage-index-shopper-system | 1 | 1 | breakage-index.offer_ambiguity.parser.v1@1 | store_domain, store_name, promo_context | offer_text, discount_value, eligible_scope, promo_code, expiration_claim, source_quote | For {{store_name}} ({{store_domain}}), describe any active promotion in {{promo_context}}. Include the discount, eligible products, promo code, and expiration only when visible. Say when terms are unclear. |
| offer_ambiguity.perplexity.v1 | offer_ambiguity | perplexity | perplexity | sonar-pro | breakage-index-shopper-system | 1 | 1 | breakage-index.offer_ambiguity.parser.v1@1 | store_domain, store_name, promo_context | offer_text, discount_value, eligible_scope, promo_code, expiration_claim, source_quote | For {{store_name}} ({{store_domain}}), describe any active promotion in {{promo_context}}. Include the discount, eligible products, promo code, and expiration only when visible. Say when terms are unclear. |
| offer_ambiguity.claude.v1 | offer_ambiguity | claude | anthropic | claude-sonnet-4-6 | breakage-index-shopper-system | 1 | 1 | breakage-index.offer_ambiguity.parser.v1@1 | store_domain, store_name, promo_context | offer_text, discount_value, eligible_scope, promo_code, expiration_claim, source_quote | For {{store_name}} ({{store_domain}}), describe any active promotion in {{promo_context}}. Include the discount, eligible products, promo code, and expiration only when visible. Say when terms are unclear. |
| offer_ambiguity.gemini.v1 | offer_ambiguity | gemini | gemini-2.5-flash | breakage-index-shopper-system | 1 | 1 | breakage-index.offer_ambiguity.parser.v1@1 | store_domain, store_name, promo_context | offer_text, discount_value, eligible_scope, promo_code, expiration_claim, source_quote | For {{store_name}} ({{store_domain}}), describe any active promotion in {{promo_context}}. Include the discount, eligible products, promo code, and expiration only when visible. Say when terms are unclear. | |
| non_public_product_exposure.chatgpt.v1 | non_public_product_exposure | chatgpt | openai | gpt-4o-mini | breakage-index-shopper-system | 1 | 1 | breakage-index.non_public_product_exposure.parser.v1@1 | store_domain, store_name, product_titles, product_urls | recommended_product_name, recommended_product_url, access_signal, source_quote | For {{store_name}} ({{store_domain}}), recommend relevant products from this candidate set only: {{product_titles}} / {{product_urls}}. If a product page appears blocked, noindexed, B2B-only, removed, or otherwise non-public, call that out explicitly. |
| non_public_product_exposure.perplexity.v1 | non_public_product_exposure | perplexity | perplexity | sonar-pro | breakage-index-shopper-system | 1 | 1 | breakage-index.non_public_product_exposure.parser.v1@1 | store_domain, store_name, product_titles, product_urls | recommended_product_name, recommended_product_url, access_signal, source_quote | For {{store_name}} ({{store_domain}}), recommend relevant products from this candidate set only: {{product_titles}} / {{product_urls}}. If a product page appears blocked, noindexed, B2B-only, removed, or otherwise non-public, call that out explicitly. |
| non_public_product_exposure.claude.v1 | non_public_product_exposure | claude | anthropic | claude-sonnet-4-6 | breakage-index-shopper-system | 1 | 1 | breakage-index.non_public_product_exposure.parser.v1@1 | store_domain, store_name, product_titles, product_urls | recommended_product_name, recommended_product_url, access_signal, source_quote | For {{store_name}} ({{store_domain}}), recommend relevant products from this candidate set only: {{product_titles}} / {{product_urls}}. If a product page appears blocked, noindexed, B2B-only, removed, or otherwise non-public, call that out explicitly. |
| non_public_product_exposure.gemini.v1 | non_public_product_exposure | gemini | gemini-2.5-flash | breakage-index-shopper-system | 1 | 1 | breakage-index.non_public_product_exposure.parser.v1@1 | store_domain, store_name, product_titles, product_urls | recommended_product_name, recommended_product_url, access_signal, source_quote | For {{store_name}} ({{store_domain}}), recommend relevant products from this candidate set only: {{product_titles}} / {{product_urls}}. If a product page appears blocked, noindexed, B2B-only, removed, or otherwise non-public, call that out explicitly. | |
| checkout_exception_risk.chatgpt.v1 | checkout_exception_risk | chatgpt | openai | gpt-4o-mini | breakage-index-shopper-system | 1 | 1 | breakage-index.checkout_exception_risk.parser.v1@1 | store_domain, store_name, checkout_country, policy_urls | payment_methods, guest_checkout_claim, shipping_country_claim, source_quote | For {{store_name}} ({{store_domain}}), tell a shopper in {{checkout_country}} which payment methods, guest checkout path, and shipping eligibility are supported. Use only public checkout, payment, help, or policy evidence from {{policy_urls}}. |
| checkout_exception_risk.perplexity.v1 | checkout_exception_risk | perplexity | perplexity | sonar-pro | breakage-index-shopper-system | 1 | 1 | breakage-index.checkout_exception_risk.parser.v1@1 | store_domain, store_name, checkout_country, policy_urls | payment_methods, guest_checkout_claim, shipping_country_claim, source_quote | For {{store_name}} ({{store_domain}}), tell a shopper in {{checkout_country}} which payment methods, guest checkout path, and shipping eligibility are supported. Use only public checkout, payment, help, or policy evidence from {{policy_urls}}. |
| checkout_exception_risk.claude.v1 | checkout_exception_risk | claude | anthropic | claude-sonnet-4-6 | breakage-index-shopper-system | 1 | 1 | breakage-index.checkout_exception_risk.parser.v1@1 | store_domain, store_name, checkout_country, policy_urls | payment_methods, guest_checkout_claim, shipping_country_claim, source_quote | For {{store_name}} ({{store_domain}}), tell a shopper in {{checkout_country}} which payment methods, guest checkout path, and shipping eligibility are supported. Use only public checkout, payment, help, or policy evidence from {{policy_urls}}. |
| checkout_exception_risk.gemini.v1 | checkout_exception_risk | gemini | gemini-2.5-flash | breakage-index-shopper-system | 1 | 1 | breakage-index.checkout_exception_risk.parser.v1@1 | store_domain, store_name, checkout_country, policy_urls | payment_methods, guest_checkout_claim, shipping_country_claim, source_quote | For {{store_name}} ({{store_domain}}), tell a shopper in {{checkout_country}} which payment methods, guest checkout path, and shipping eligibility are supported. Use only public checkout, payment, help, or policy evidence from {{policy_urls}}. | |
| competitor_in_branded_prompt.chatgpt.v1 | competitor_in_branded_prompt | chatgpt | openai | gpt-4o-mini | breakage-index-shopper-system | 1 | 1 | breakage-index.competitor_in_branded_prompt.parser.v1@1 | store_domain, store_name, primary_category, competitor_set | brand_answered, competitor_mentions, competitor_recommendations, source_quote | A shopper asks specifically about {{store_name}} for {{primary_category}}. Answer the branded question directly. If you mention or recommend alternatives from {{competitor_set}}, label whether they are neutral mentions or explicit recommendations over {{store_name}}. |
| competitor_in_branded_prompt.perplexity.v1 | competitor_in_branded_prompt | perplexity | perplexity | sonar-pro | breakage-index-shopper-system | 1 | 1 | breakage-index.competitor_in_branded_prompt.parser.v1@1 | store_domain, store_name, primary_category, competitor_set | brand_answered, competitor_mentions, competitor_recommendations, source_quote | A shopper asks specifically about {{store_name}} for {{primary_category}}. Answer the branded question directly. If you mention or recommend alternatives from {{competitor_set}}, label whether they are neutral mentions or explicit recommendations over {{store_name}}. |
| competitor_in_branded_prompt.claude.v1 | competitor_in_branded_prompt | claude | anthropic | claude-sonnet-4-6 | breakage-index-shopper-system | 1 | 1 | breakage-index.competitor_in_branded_prompt.parser.v1@1 | store_domain, store_name, primary_category, competitor_set | brand_answered, competitor_mentions, competitor_recommendations, source_quote | A shopper asks specifically about {{store_name}} for {{primary_category}}. Answer the branded question directly. If you mention or recommend alternatives from {{competitor_set}}, label whether they are neutral mentions or explicit recommendations over {{store_name}}. |
| competitor_in_branded_prompt.gemini.v1 | competitor_in_branded_prompt | gemini | gemini-2.5-flash | breakage-index-shopper-system | 1 | 1 | breakage-index.competitor_in_branded_prompt.parser.v1@1 | store_domain, store_name, primary_category, competitor_set | brand_answered, competitor_mentions, competitor_recommendations, source_quote | A shopper asks specifically about {{store_name}} for {{primary_category}}. Answer the branded question directly. If you mention or recommend alternatives from {{competitor_set}}, label whether they are neutral mentions or explicit recommendations over {{store_name}}. |
Sample Size
The External Top-500 floor is a fast-launch compromise: large enough for a credible inaugural benchmark, with uncertainty disclosed for every headline and agent cut.
Per-agent cuts can publish only with their own denominator and uncertainty disclosure; they do not replace the all-store headline denominator.
For the Summer 2026 report, per-agent cuts are filterable data only. They are not used for launch copy, vendor rankings, or agent-superlative claims.
Summer 2026 public artifacts report all-store cuts only. Per-vertical cuts were withheld because the vertical validation review did not pass the publication threshold.
Corpora And Exclusions
The External Top-500 corpus uses Tranco list VNGQN from 2026-05-21, filtered to retail and Shopify-detected domains. Included stores: 500. The internal corpus is disclosed separately and is never blended into the External Top-500 headline number. Canonical path: /resources/research/agentic-commerce-breakage-index.
The internal StoreSteady scan corpus is self-selected and reported with a self-selection disclosure when available.
- homepage_crawl_failed: Homepage did not resolve during the report window
- policy_pages_missing: Required policy pages were not observable
- fewer_than_10_indexable_pdps: Store had fewer than 10 indexable product pages
- aggressive_antibot: Store blocked the scanner with aggressive anti-bot controls
- pdps_behind_login: Product pages were behind login
- no_observable_product_truth: Product truth was not observable for this metric
- no_observable_policy_truth: Policy truth was not observable for this metric
- no_observable_offer_truth: Offer truth was not observable for this metric
- no_observable_non_public_signal: No externally verifiable non-public signal was observable
- no_observable_checkout_truth: Checkout capability truth was not observable
- unsupported_agent_surface: Agent surface is not supported by the frozen protocol
- merchant_opted_out: Merchant opted out of internal aggregate reporting
- calibration_ineligible: Candidate did not pass calibration eligibility
- reproducibility_failed: Candidate did not survive the 2-of-3 reproducibility rule
Denominator Rules
The headline unit is the store, not the finding and not the response.
Each metric publishes its own evaluable_store_n. Stores with no observable truth for a metric are excluded from that metric denominator and counted in the exclusion funnel.
The at-least-one-breakage headline uses stores evaluable for the full metric bundle, or explicitly discloses the minimum metric-evaluable threshold used.
Crawler blocks, robots restrictions, password or challenge gates, and other non-circumvented access controls are counted separately as access_friction_observed. They are not counted as product, policy, offer, checkout, hidden-product, or competitor breakage unless the relevant detector also has enough observable evidence.
| metric | corpus | vertical | agent | evaluable_store_n | stores_with_confirmed_breakage | excluded_store_n |
|---|---|---|---|---|---|---|
| checkout_exception_risk | external_top_500 | all | all | 104 | 26 | 0 |
| checkout_exception_risk | external_top_500 | all | chatgpt | 19 | 0 | 0 |
| checkout_exception_risk | external_top_500 | all | claude | 104 | 1 | 0 |
| checkout_exception_risk | external_top_500 | all | gemini | 33 | 1 | 0 |
| checkout_exception_risk | external_top_500 | all | perplexity | 97 | 24 | 0 |
| non_public_product_exposure | external_top_500 | all | all | 272 | 5 | 0 |
| non_public_product_exposure | external_top_500 | all | chatgpt | 271 | 4 | 0 |
| non_public_product_exposure | external_top_500 | all | claude | 268 | 4 | 0 |
| non_public_product_exposure | external_top_500 | all | gemini | 193 | 3 | 0 |
| non_public_product_exposure | external_top_500 | all | perplexity | 256 | 2 | 0 |
| product_truth_mismatch | external_top_500 | all | all | 273 | 12 | 0 |
| product_truth_mismatch | external_top_500 | all | chatgpt | 31 | 0 | 0 |
| product_truth_mismatch | external_top_500 | all | claude | 273 | 2 | 0 |
| product_truth_mismatch | external_top_500 | all | gemini | 272 | 9 | 0 |
| product_truth_mismatch | external_top_500 | all | perplexity | 121 | 3 | 0 |
Agent Access Friction
StoreSteady does not bypass anti-bot systems, login walls, payment gates, robots restrictions, or challenge pages during this benchmark.
A store can remain in the research frame while a specific detector job is not evaluable. The scan-limitations table separates access_friction_observed from not_evaluable_for_detector so readiness signals do not inflate breakage rates.
Access friction is merchant-actionable product value for Pro Scan because agent-like systems may fail to read, compare, or recommend a store even when no detector-specific mismatch can be evaluated.
| classification | access_friction_kind | status | reason_code | scanned_store_n | affected_store_n | store_percentage | affected_job_n |
|---|---|---|---|---|---|---|---|
| access_friction_observed | all | all_access_friction | 500 | 210 | 0.42 | 5040 | |
| access_friction_observed | crawl_blocked | failed | crawl_blocked_by_robots_without_classified_pages | 500 | 2 | 0.004 | 48 |
| access_friction_observed | crawl_blocked | failed | crawl_failed_without_classified_pages | 500 | 208 | 0.416 | 4992 |
| not_evaluable_for_detector | all | all_not_evaluable_for_detector | 500 | 290 | 0.58 | 2941 | |
| not_evaluable_for_detector | unevaluable | inconclusive_checkout | 500 | 6 | 0.012 | 18 | |
| not_evaluable_for_detector | unevaluable | missing_product_claim | 500 | 265 | 0.53 | 485 | |
| not_evaluable_for_detector | unevaluable | no_matching_truth_product | 500 | 7 | 0.014 | 7 | |
| not_evaluable_for_detector | unevaluable | no_observable_checkout_truth | 500 | 180 | 0.36 | 701 | |
| not_evaluable_for_detector | unevaluable | no_observable_non_public_signal | 500 | 17 | 0.034 | 68 | |
| not_evaluable_for_detector | unevaluable | no_observable_offer_truth | 500 | 160 | 0.32 | 640 | |
| not_evaluable_for_detector | unevaluable | no_observable_policy_truth | 500 | 178 | 0.356 | 712 | |
| not_evaluable_for_detector | unevaluable | no_observable_product_truth | 500 | 17 | 0.034 | 68 | |
| not_evaluable_for_detector | unevaluable | unknown_product_status | 500 | 2 | 0.004 | 7 | |
| not_evaluable_for_detector | unevaluable | unparseable_checkout_response | 500 | 108 | 0.216 | 188 | |
| not_evaluable_for_detector | unevaluable | unparseable_offer_response | 500 | 21 | 0.042 | 47 |
Tolerance Bands
Each metric uses the same public tolerance bands throughout the Summer 2026 report.
- Product truth mismatch: Price differs by more than 2% or more than $5, stock contradicts the PDP, product does not resolve to a public PDP, or material attributes contradict product evidence. Attribute claims use the two-judge plus human-tiebreak calibration path.
- Return or shipping policy misstatement: Numeric return, fee, and shipping terms must match the policy page exactly. Categorical claims are binary. Free-text policy claims use the two-judge plus human-tiebreak calibration path.
- Offer ambiguity: Wrong code, expired offer, or overstated discount scope counts as breakage when public offer or safe dry-run evidence contradicts the agent answer. Ambiguous offer text is held for calibration review when parser output is not deterministic.
- Non-public product exposure: Only externally verifiable signals count: 401/403, 410, noindex, robots disallow, or explicit B2B-only copy. Public absence alone is not proof. Draft/private status requires connected-store evidence and is not inferred for the external corpus.
- Checkout exception risk: Unsupported payment method, guest-checkout mismatch, or shipping-country mismatch counts only when public checkout, payment, help, or policy evidence contradicts the agent answer. Checkout dry-runs are non-transactional and blocked runs are exclusions, not breakage.
- Competitor in branded prompt: Neutral competitor mentions are tracked separately. Explicit recommendations over the named merchant count as breakage. The prompt must explicitly contain the merchant brand for this metric.
Calibration And False-Positive Handling
The calibration set samples flagged candidates and non-finding candidates. Public calibration rows expose labels and rubric versions only, not raw store names, raw URLs, or raw AI responses.
Two independent LLM judges score each candidate. Disagreements require human tiebreak. Krippendorff alpha is published per metric.
Headline metrics require alpha at least 0.80 and calibration accuracy at least 95%. Alpha from 0.667 to 0.80 is tentative. Alpha below 0.667 is dropped.
Public evidence examples are derived from normalized calibration evidence packages. A fail-closed safety review blocks any example containing store names, merchant domains, raw URLs, source quotes, raw AI responses, exact promo codes, or merchant-identifying details.
Judge 1 model: gpt-4.1. Judge 2 model: claude-sonnet-4-6.
Rubric versions: product_truth_mismatch: product_attribute_v1, product_attribute_v2, policy_misstatement: policy_free_text_v1, policy_free_text_v2, offer_ambiguity: offer_ambiguity_v1, offer_ambiguity_v2, non_public_product_exposure: non_public_product_exposure_v1, non_public_product_exposure_v2, checkout_exception_risk: checkout_exception_risk_v1, checkout_exception_risk_v2, competitor_in_branded_prompt: branded_competitor_v1, branded_competitor_v2.
Two independent LLM judges label calibration rows; disagreements are resolved using the human-authored labeling policy and persisted as human tiebreak labels.
| metric | tier | judge_alpha | calibration_accuracy | calibration_flagged_n | calibration_nonfinding_n | false_negative_rate |
|---|---|---|---|---|---|---|
| product_truth_mismatch | headline | 0.9233218943033631 | 99.0% | 61 | 139 | 1.4% |
| policy_misstatement | dropped | 0.2703367385546728 | 56.6% | 100 | 100 | 0.0% |
| offer_ambiguity | dropped | 0.18794123258920037 | 99.0% | 5 | 195 | 0.5% |
| non_public_product_exposure | headline | 1 | 100.0% | 16 | 184 | 0.0% |
| checkout_exception_risk | tentative | 0.7939404372525392 | 100.0% | 100 | 100 | 0.0% |
| competitor_in_branded_prompt | dropped | 0.982918789331735 | 67.5% | 100 | 100 | 0.0% |
Limitations
The Breakage Index measures answer fidelity and access friction. It does not measure conversion impact, revenue loss, refund rate, chargeback rate, or support-ticket volume.
This report reflects stores that were evaluable inside the report window. Each metric publishes its own evaluable_store_n, and stores with missing truth are counted in the exclusion funnel instead of being treated as breakage.
Stores that block crawler or agent-like access are retained in the research frame when they are otherwise eligible retail Shopify storefronts. Blocked access is reported separately as access friction, not as a detector mismatch.
The internal corpus consists of stores that voluntarily ran a StoreSteady scan during the reporting period. It over-represents operators who already suspected AI-search issues and is never blended into the External Top-500 headline number.
The External Top-500 is drawn from a frozen Tranco snapshot, filtered to retail and Shopify-detected domains. Long-tail, non-English, and non-Shopify stores are under-represented by design and disclosed as out of scope for this inaugural index.
- No per-store names, raw URLs, or raw AI responses are published in public artifacts.
- External non-public product evidence is limited to public signals. Shopify draft status is counted only when connected-store evidence proves it.
- Checkout and offer probes are non-transactional. Store blocks, anti-bot walls, and inconclusive dry-runs are access-friction or limitation signals, not breakage.
- Metrics below 95% calibration accuracy are not published in this report. Metrics with judge alpha from 0.667 to 0.80 and calibration accuracy at least 95% are tentative. Headline metrics require alpha at least 0.80 and calibration accuracy at least 95%.
- Future quarterly runs depend on distribution pickup and detector reuse in Pro Scan, not backlinks alone.
Correction Policy
Methodology gaps, reproducible errors, or source corrections can be sent to research@storesteady.com. Corrections are published with the version history below.
Sources
- Tranco list - External Top-500 corpus snapshot source.
- Tranco research paper - Background on manipulation resistance for the ranking source.
- Google Merchant Center product data specification - Product-data accuracy context; not used as a StoreSteady impact claim.
- Shopify product taxonomy - Shopify product-data context for category and product detail fields.
- OpenAI Shopping with ChatGPT Search - Agentic shopping context for product answers and purchase links.
- OpenAI crawlers - Crawler and access context for AI systems.
- Google robots.txt specification - Crawler access-control context; not counted as breakage by itself.
- Google robots meta tag documentation - Indexing and noindex context for non-public product signals.
- Shopify robots.txt editing - Shopify storefront crawler-control context.
- Baymard checkout usability benchmark - Checkout-friction context; not used to estimate StoreSteady conversion impact.
- Baymard cart-abandonment guidance - Checkout-friction context; not used as a causal revenue claim.
- NRF and Happy Returns 2024 retail returns report - Returns-market context; not used to estimate StoreSteady refund impact.
- FTC Advertising FAQs - Offer and pricing claim context; not used as a StoreSteady legal finding.
- Cochran sample-size formula - Sample-size floor used for single-metric estimates.
- Bonferroni correction reference - Reference for multi-metric uncertainty disclosure on the external corpus.
- Krippendorff's alpha thresholds - Reliability threshold framing for judge agreement.
- MLPerf LLM inference methodology - Reference pattern for public benchmark calibration disclosure.
- LLM-as-judge transparency paper - Reference for publishing judge prompts, model IDs, rubrics, and limits.
- Pew margin-of-error disclosure pattern - Public disclosure pattern for uncertainty and subgroup caveats.
- Stack Overflow survey methodology - Self-selection disclosure pattern for the internal StoreSteady corpus.
- Cloudflare Radar methodology pattern - Public methodology pattern for source limits and transparent caveats.
- Edelman Trust Barometer methodology - Public trust-report methodology reference.
- Profound Index - Visibility-focused competitor benchmark reference.
- Otterly e-commerce ranking - Visibility-focused competitor ranking reference.
- Bain agentic-commerce forecast - Market-context source; not used as a detector claim.
- McKinsey shopping in the age of AI - Market-context source; not used as a detector claim.
- BrightEdge AI search insight - AEO and citation-distribution context for publication planning.
- Evidently AI LLM evaluation metrics - Tolerance-band and evaluation framing reference.
Version History
- 2026-05-22: 2026-summer.c4-v1 - Initial C4 methodology content model and renderer contract.
- 2026-05-24: 2026-summer.public-v1 - Finalized public Summer 2026 release framing, added denominator-forward funnel disclosure, and blocked relisting on complete provenance.