GrowthOS

Agent Task Completion Index

How often agents complete standardized real-world tasks across software products.

Comparable outcomes only

Production behavior and controlled behavior stay distinct.

Rows are grouped and ranked only when the standardized task, metric kind, methodology version, and time-window policy are comparable.

PRODUCTION

Production ATCR

Completed reconstructed production tasks divided by eligible reconstructed production tasks. This measures what agents actually completed.

CONTROLLED

Eval pass rate

Controlled runs meeting the stated success criterion divided by eligible controlled runs. This is not relabeled as production performance.

PUBLICATION GATE

Every row must disclose

  • Standardized task
  • Completion criterion
  • Production or Eval provenance
  • Attempts and sample size
  • Agent and model context
  • Time window
  • Evidence-confidence treatment
  • Methodology version
  • Limitations and permission

Current publication state

The Index begins with evidence, not placeholder ranks.

No public record currently passes the full publication gate. We will not fill the space with fabricated products, hidden samples, or incomparable scores.

Evidence already inspectable

Start with the task studies.

Study pages show the evidence structure before enough comparable records exist for an Index cohort.