Blog
Why we benchmark stores, not models
09.01.26Model benchmarks answer the wrong question for merchants. We benchmark the other half of the transaction: the store.
What 500 Shopify stores look like to a machine
09.02.26Assembling the benchmark panel, before any agent ran a single task. Preliminary observations.
The anatomy of a failed checkout
09.03.26One representative failure, traced end to end. The journey was 90 percent complete.
Agent readiness is not a vibe
09.04.26A store does not feel agent-ready. It scores agent-ready. How the scorecard stays honest.
Notes on pricing: extraction is not understanding
09.05.26Read the price sounds trivial. It splits into three abilities, and the hardest one is not extraction.
The machine customer already exists
09.08.26Agents are already shopping at small scale, badly. Every failed attempt is a sale that silently did not happen.
