What gets measured
A session is scored against whatever its scenario defines in Defining Good:
Depending on the scenario, scoring covers the conversation, the outputs, or both. A decision task is scored on its decision and rationale; a conversation also on what was elicited and how; a journey also on what held up across sessions.
Reading a report
A scored session shows each criterion with its result and the reasoning behind it, grounded in the transcript and outputs. You can see not just that a failure metric fired, but the moment in the conversation it points to. The score is an entry point into the session, not a verdict that replaces it.Comparing results
Because criteria are fixed per scenario, results aggregate cleanly:- Across operators - a cohort shows per-operator results on the same scenarios, so patterns separate from individuals
- Across agents - a benchmark run shows the same breakdown for each agent, prompt version, or model configuration you test
- People against AI - the same scenario, the same criteria, side by side
Treat a single session as one observation, not a conclusion. The platform makes it cheap to run a scenario several times; spreads across repeated sessions tell you what is signal and what is one good or bad day.