Machine Commerce
Machine Commerce
Agent Readiness Score: Methodology

Agent Readiness Score: Methodology

By Machine Commerce September 2026

How a store's score out of 100 is computed, what the letter grades mean, and the rerun rule that keeps scores honest.

The Score

Each store receives a score from 0 to 100. The score is the weighted sum of task results across the six benchmark tasks. Weights reflect task difficulty and commercial significance. Checkout Completion carries the highest weight.

The Checks

Each task produces one or more binary checks. A check is a specific observable outcome: the agent read the correct price, the agent added the correct variant, the agent reached the payment confirmation screen. Partial credit is not awarded within a check. Either it passed or it did not.

The Grade

Scores map to letter grades: A (90-100), B (75-89), C (60-74), D (45-59), F (below 45). Grades are intended as a fast summary for non-technical stakeholders. The underlying score and check list are always published alongside the grade.

Reproducibility

Every check must produce the same result across three independent runs before it is recorded. A check that produces mixed results is flagged as unstable and excluded from scoring until the instability is resolved. Unstable checks are reported transparently in the published scorecard.

Limitations

Scores reflect agent performance on a specific task suite at a specific point in time. A high score does not guarantee that every possible agent can complete every possible purchase in a given store. Stores change. Agents change. We re-run the full panel every quarter.