Research
Machine Commerce publishes benchmark data, methodology, and findings from running AI shopping agents against real Shopify stores. The research covers the full agentic commerce workflow, from product discovery to checkout completion, and is designed to be reproducible, openly cited, and honest about failure modes.
Publications
The Machine Commerce Benchmark: Task Suite v0.1
09.01.26The six commerce tasks every benchmarked agent attempts inside real stores, and the procedure that keeps runs comparable.
Agent Readiness Score: Methodology
09.02.26How a store's score out of 100 is computed, what the letter grades mean, and the rerun rule that keeps scores honest.
A Failure Taxonomy for Machine-Mediated Shopping
09.03.26Seven classes cover every failure we have logged so far. The taxonomy is versioned and open.
Dataset Note: The 500-Store Panel
09.04.26Composition, selection criteria, refresh cadence, and run conduct for the benchmark's store panel.






