
Agent Readiness Score: Methodology
How a store's score out of 100 is computed, what the letter grades mean, and the rerun rule that keeps scores honest.
The Score
Each store receives a score from 0 to 100. The score is the weighted sum of task results across the six benchmark tasks. Weights reflect task difficulty and commercial significance. Checkout Completion carries the highest weight.
The Checks
Each task produces one or more binary checks. A check is a specific observable outcome: the agent read the correct price, the agent added the correct variant, the agent reached the payment confirmation screen. Partial credit is not awarded within a check. Either it passed or it did not.
The Grade
Scores map to letter grades: A (90-100), B (75-89), C (60-74), D (45-59), F (below 45). Grades are intended as a fast summary for non-technical stakeholders. The underlying score and check list are always published alongside the grade.
Reproducibility
Every check must produce the same result across three independent runs before it is recorded. A check that produces mixed results is flagged as unstable and excluded from scoring until the instability is resolved. Unstable checks are reported transparently in the published scorecard.
Limitations
Scores reflect agent performance on a specific task suite at a specific point in time. A high score does not guarantee that every possible agent can complete every possible purchase in a given store. Stores change. Agents change. We re-run the full panel every quarter.



