
Why we benchmark stores, not models
Every week, a new benchmark tells us which model is smartest. Almost none of them answer the question a merchant actually has: can an agent complete a purchase in my store?
The model is only half of the transaction. The other half is the store: its product pages, variant pickers, cart logic, checkout flow, and pricing behavior. A capable agent can still fail inside a store that was built assuming eyes and thumbs.
So we benchmark the other half. Machine Commerce runs AI shopping agents against real e-commerce stores and measures where the journey breaks: discovery, selection, cart, checkout, and everything after.
This site is where we publish the results. The benchmark is public, the scores are public, and the failures are documented. When the numbers are wrong, we say so and rerun.
Commerce is getting a second customer. We intend to measure what that customer can and cannot do.





