Production ATCR
Completed reconstructed production tasks ÷ eligible reconstructed production tasks.
Answers what agents actually completed in your live product.GrowthOS reconstructs how agents discover, choose, and use your product in production, then runs the same real-world tasks across any agent and model.
Engine 01 · Production
One agent objective can cross docs, authentication, an API, an SDK, and a backend, including through parallel sessions and delegated work. GrowthOS joins the evidence into a journey graph without inventing certainty.
Every connection carries its confidence, supporting evidence, and unresolved alternatives.
How reconstruction worksTASK · SDK-001
The agent finds the SDK overview through documentation search and opens the installation guide.
Engine 02 · Controlled
Turn a production journey or qualified workflow into a managed Eval. Hold the success criterion steady while you compare agents, models, interfaces, documentation, and product treatments.
Completion is only one answer. Inspect friction, latency, tokens, retries, failure modes, and the trajectory behind the result.
Explore GrowthOS Evals| Agent | Model | Treatment | Stage | Completion | Friction | Latency | Tokens | Retries | Failure mode | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|
| Agent A | Model 1 | Baseline | configure | Blocked | 4/5 | 41s | 8,650 | 2 | Auth scope unresolved | |
| Agent A | Model 1 | Docs treatment | complete | Completed | 1/5 | 29s | 6,120 | 0 | - | |
| Agent B | Model 2 | Interface treatment | complete | Completed | 2/5 | 33s | 7,010 | 1 | - | |
| Agent C | Model 3 | Docs treatment | complete | Completed | 1/5 | 31s | 6,550 | 0 | - |
Completion uses the task criterion. Friction counts avoidable recovery work. Latency, tokens, and retries describe the controlled run. Failure mode names the observed blocker.
One evidence system
The two engines share the same task definition and evidence model, creating a loop from real behavior to controlled improvement and back to production verification.
The outcome metric
Agent Task Completion Rate (ATCR) gives production behavior an outcome unit. It belongs beside, but never collapses into, controlled Eval pass rate or reconstruction confidence.
Completed reconstructed production tasks ÷ eligible reconstructed production tasks.
Answers what agents actually completed in your live product.Controlled runs that meet a fixed success criterion ÷ eligible controlled runs.
Answers what completed under comparable Eval conditions.The strength of the evidence connecting activity into a task journey.
Qualifies the reconstruction; it is not a performance score.The full agent experience
Use one connected evidence view to see whether agents can discover, access, use, and transact with your product, not four disconnected scorecards.
ACTIVE LAYER · DISCOVERY
A lower-commitment starting point
See why sessions and pageviews miss agent demand, how to define production task completion, and which evidence makes agent behavior actionable.
Request the report01 · Invisible demand
02 · What to measure
ATCR
Start with one task
We’ll scope the surfaces, evidence, completion criterion, and the clearest first reconstruction.
Not ready for a reconstruction? Request the Agent Traffic Blindspot Report.