Agent Economics

Traditional SaaS economics often starts with “users × monthly price” on the revenue side and servers, bandwidth, and storage on the cost side. Agent products do not fit that shape. Users consume inference, tool calls, repeated context, retries, and human handoffs rather than just seats.

The goal of agent economics is not to memorize today’s model prices. It is to answer three durable questions:

  1. Why does one task cost what it costs?
  2. Which costs grow with task length?
  3. How do cost controls and business value fit into the same model?

Where Cost Comes From

One agent task usually contains six cost categories:

CostMeaningCommon controls
Input tokensSystem prompt, user input, history, and tool results entering contextPrompt caching, trimming, retrieval granularity
Cached tokensStatic prefixes read from cache at lower pricesKeep system prompts and tool descriptions stable
Output tokensNatural language, plans, explanations, or tool arguments generated by the modelReduce unnecessary reasoning output; prefer direct tool calls
Tool / runtimeBrowser, sandbox, database, external API, file system, or session infrastructureTool limits, batching, sandbox lifecycle management
Retry / failureExtra inference, rollback, and redo work after wrong pathsEvals, budget stops, user confirmation points
Human reviewReview, takeover, correction, and process trainingPlace HITL only at high-risk points

Tokens are the easiest part to measure, but they are not the whole cost. A seemingly cheap agent can still have poor economics if it fails often and requires humans to clean up the work.

Three Dimensions Of A Token

Each token consumes three resources at once:

ResourceMeaningBusiness impact
MoneyToken count × current model priceBill and gross margin
LatencyInput processing and output generation take timeUser-facing wait time
CapacityTokens occupy the context windowTask depth and retained state

Before optimizing, name the target. Many choices are trade-offs rather than pure improvements:

  • Compacting history can reduce future input, but compaction itself costs tokens and may lose detail.
  • A stronger model costs more per token, but may take fewer steps and retry less.
  • A cheaper model lowers per-step cost, but may increase step count, failure rate, and human takeover.
  • A longer context can reduce handoff overhead, but may also make every step carry more history.

Agent cost work is therefore not just “save tokens.” It is deciding which tokens are worth spending and which tasks should not be delegated to an agent.

Three Governing Principles

Keep static content cacheable. If the system prompt, tool descriptions, or policy blocks change on every step, prompt caching breaks. Preserving a large static prefix usually matters more than manually deleting a few hundred tokens.

Long-running history compounds. Multi-step agents often resend prior messages and tool results into context. Without trimming, compaction, or externalized state, history cost grows quickly with step count.

Failure is rarely cheaper than success. Tokens, tool calls, and human time spent on a failed path are sunk. A late-stage failure in a long task can cost more to repair than one successful run.

A Reference Bill

The following example is for intuition, not a quote. Suppose a user asks an agent to triage the latest 20 emails by project, and the task runs for 12 steps. Using one public pricing snapshot and synthetic token usage, the bill is approximately $0.19.

The rough distribution is:

  • 41% Conversation history: prior messages and tool results re-enter later inputs.
  • 23% Model output: responses, plans, and tool arguments.
  • 22% Tool results: returned tool content that later becomes context.
  • 13% System prompt + tool descriptions: static prefix, assuming strong cache hits.
  • < 1% Sandbox + storage: small in this example, but workload-dependent in real systems.

This distribution is not a law. It shows that in medium-length, multi-step agent tasks with history replay, history, tool results, and output often deserve more attention than startup prompt overhead.

Section Guide

SectionWhat it coversAudience
Cost ModelFormulas and a synthetic example for input, cache, output, tools, failures, and human reviewEngineering, product, finance
Controls and ROIBudgets, model routing, caching, circuit breakers, monitoring, and conservative cost-benefit estimationOperations, sales, procurement

If you read only one page, start with Cost Model. It turns “agents are expensive” into fields that can be measured.

Was this page helpful?