Metrics
Agent product metrics need to move from “was the product used?” to “was work completed?” Opening the app, starting a task, and consuming tokens do not necessarily mean value was created. The useful questions are: how many tasks did the agent do, how many succeeded, how good were the results, and did they produce business value?
Why Usage Metrics Are Not Enough
MAU, DAU, session length, and message count still matter, but they only show that users touched the product. They do not show whether the agent delivered work.
A user starting 5 tasks a day and failing every time looks active, but may be burning cost, creating frustration, and increasing churn risk. Another user may run one task a week, but complete an important workflow every time. The second user can be more valuable.
The basic measurement unit for an agent product should therefore be the task, not the message, session, or app open.
Three Metric Layers
Volume
Volume shows how much work the agent is asked to do.
| Metric | Meaning |
|---|---|
| Tasks started | How many executable requests users submit |
| Tasks completed | How many tasks the agent completes |
| Tasks per user | Demand penetration and usage depth |
| Task type mix | Which workflows users delegate most often |
Volume is the base layer, but not a success signal by itself. A rise in task starts can mean growing demand, or it can mean users are retrying the same failed job.
Quality
Quality shows whether the agent is reliable.
| Metric | Meaning |
|---|---|
| Success rate | Share of started tasks that reach the goal |
| First-pass success rate | Share that succeed without retry or human fallback |
| HITL rate | Share requiring human confirmation, review, or takeover |
| Recovery rate | Share rescued by retry, fallback, or human takeover |
| Rework rate | Share whose output needs human or user correction |
Quality metrics must be tied to task type. Email triage, code edits, customer refunds, and data analysis have different success criteria. One global success rate cannot explain all failures.
Value
Value shows whether the agent is worth running.
| Metric | Meaning |
|---|---|
| Cost per successful task | Cost divided by successful tasks, not all starts |
| Hours saved | Human time replaced or shortened by the agent |
| Net value per task | Human baseline value minus agent cost, review cost, and failure adjustment |
| Customer-visible ROI | A value measure the customer can understand and audit |
| Gross-margin contribution | Revenue minus inference, tool, runtime, and human costs |
The value layer depends on Agent Economics: cost is not only tokens, but also tools, runtime, retries, and human review.
How The Layers Fit
The layers cannot substitute for each other:
- Tracking only volume encourages more tasks even if they fail.
- Tracking only quality can make the team overly conservative and shrink the automation surface.
- Tracking only value can hide early learning signals and usage shifts.
A practical dashboard often follows this chain:
task volume -> success rate -> cost per successful task -> net value per task -> retention / expansion
First confirm demand exists. Then confirm the agent can complete work reliably. Then confirm each successful result has acceptable cost. Only then does long-term revenue become meaningful.
Section Guide
| Section | Question |
|---|---|
| North-star Metric | Which metric should be the team’s primary goal under each business model? |
| Unit Economics | Does each user, task, or contract actually make money? Do heavy users consume margin? |
Boundary With Other Sections
- economics defines the cost model and ROI formulas.
- metrics turns cost, quality, and business outcomes into operating metrics.
- operations turns those metrics into dashboards, alerts, and runtime controls.