GrowthOS

Agents are becoming software's primary users.

GrowthOS reconstructs how agents discover, choose, and use your product in production, then runs the same real-world tasks across any agent and model.

Trusted by teams at

Engine 01 · Production

Reconstruct the task your existing analytics split apart.

One agent objective can cross docs, authentication, an API, an SDK, and a backend, including through parallel sessions and delegated work. GrowthOS joins the evidence into a journey graph without inventing certainty.

Every connection carries its confidence, supporting evidence, and unresolved alternatives.

How reconstruction works
Illustrative reconstruction

TASK · SDK-001

Install the SDK and send the first valid event.

completed · 00:00
discoverDocsExact

Discover the SDK

The agent finds the SDK overview through documentation search and opens the installation guide.

1 / 6

Engine 02 · Controlled

Test the same real-world task across every agent and model.

Turn a production journey or qualified workflow into a managed Eval. Hold the success criterion steady while you compare agents, models, interfaces, documentation, and product treatments.

Completion is only one answer. Inspect friction, latency, tokens, retries, failure modes, and the trajectory behind the result.

Explore GrowthOS Evals
Illustrative reconstruction
Illustrative Eval comparison
AgentModelTreatmentStageCompletionFrictionLatencyTokensRetriesFailure modeEvidence
Agent AModel 1BaselineconfigureBlocked4/541s8,6502Auth scope unresolved
Agent AModel 1Docs treatmentcompleteCompleted1/529s6,1200-
Agent BModel 2Interface treatmentcompleteCompleted2/533s7,0101-
Agent CModel 3Docs treatmentcompleteCompleted1/531s6,5500-
How each metric is defined

Completion uses the task criterion. Friction counts avoidable recovery work. Latency, tokens, and retries describe the controlled run. Failure mode names the observed blocker.

One evidence system

Production tells you what to test. Evals tell you what to change.

The two engines share the same task definition and evidence model, creating a loop from real behavior to controlled improvement and back to production verification.

  1. 01

    Observe production

    Collect evidence across the surfaces agents use.

  2. 02

    Reconstruct the task

    Infer objective, path, barrier, and outcome with confidence.

  3. 03

    Run the Eval

    Replay the task under a controlled success criterion.

  4. 04

    Change the product

    Test docs, access, interface, or workflow treatments.

  5. 05

    Verify in production

    Look for the subsequent real-world outcome.

  6. 06

    Expand the task set

    Turn verified journeys into a living Eval program.

The outcome metric

Measure completed tasks, not agent traffic.

Agent Task Completion Rate (ATCR) gives production behavior an outcome unit. It belongs beside, but never collapses into, controlled Eval pass rate or reconstruction confidence.

01

Production ATCR

Completed reconstructed production tasks ÷ eligible reconstructed production tasks.

Answers what agents actually completed in your live product.
02

Eval pass rate

Controlled runs that meet a fixed success criterion ÷ eligible controlled runs.

Answers what completed under comparable Eval conditions.
03

Evidence confidence

The strength of the evidence connecting activity into a task journey.

Qualifies the reconstruction; it is not a performance score.

The full agent experience

A task succeeds only when every layer holds.

Use one connected evidence view to see whether agents can discover, access, use, and transact with your product, not four disconnected scorecards.

ACTIVE LAYER · DISCOVERY

Can the agent find the right path?

Trace how an agent discovers product pages, documentation, schemas, and integration surfaces, and whether the chosen route matches the task.
  1. 01Search and referral path
  2. 02Resource selection
  3. 03Task-fit signal

A lower-commitment starting point

Request the Agent Traffic Blindspot Report.

See why sessions and pageviews miss agent demand, how to define production task completion, and which evidence makes agent behavior actionable.

Request the report
Illustrative report spread

01 · Invisible demand

02 · What to measure

ATCR

Questions about the evidence

Start with one task

Show us one task agents should be able to complete.

We’ll scope the surfaces, evidence, completion criterion, and the clearest first reconstruction.

Request a task reconstructionTalk to a founder
Not ready for a reconstruction? Request the Agent Traffic Blindspot Report.