GrowthOS

Evals for the tasks agents actually perform.

Turn a real production journey or qualified workflow into a controlled task. Compare agents, models, documentation, interfaces, prompts, and product treatments under the same success criteria.

Comparable runs

Measure more than a final pass.

Completion

Did the run satisfy the task criterion?

Friction

Where did the run encounter avoidable effort?

Latency

How long did the task take?

Tokens

What model usage did it require?

Retries

How many attempts were needed?

Failure mode

What prevented the run from completing?

Production ↔ controlled

The same task, two kinds of evidence.

Production reconstruction discovers what matters. Evals make treatments comparable. Production verification shows what changed in the real product.

Production and Eval evidence
MetricSourceAnswersLimitation
Production ATCRReconstructed real product behaviorWhat agents actually completedVaries with available production evidence
Eval pass rateControlled runs with shared criteriaWhat comparable runs completedControlled results are not production proof

Illustrative matrix

Compare the run, not just the model.

Filter a claim-safe example by product treatment while keeping completion, friction, latency, tokens, retries, and failure mode visible together.

Illustrative reconstruction
Illustrative Eval comparison
AgentModelTreatmentStageCompletionFrictionLatencyTokensRetriesFailure modeEvidence
Agent AModel 1BaselineconfigureBlocked4/541s8,6502Auth scope unresolved
Agent AModel 1Docs treatmentcompleteCompleted1/529s6,1200-
Agent BModel 2Interface treatmentcompleteCompleted2/533s7,0101-
Agent CModel 3Docs treatmentcompleteCompleted1/531s6,5500-
How each metric is defined

Completion uses the task criterion. Friction counts avoidable recovery work. Latency, tokens, and retries describe the controlled run. Failure mode names the observed blocker.

Evaluate the task that matters.

Bring one workflow, one success criterion, and the agent experience you want to improve.