Skip to main content
Sign in
LockedIn Labs · Practitioner guide

Token economics: measure cost per accepted outcome.

Start with the work you can accept. Then account for the tokens, retries, tools, and human effort required to produce it.

By LockedIn Labs · · 8 min read

Download the scorecard · CSV
01 · The unit

Define an accepted result before counting consumption.

Token volume describes how much a system processed. It does not establish how much useful work it completed. A team considering a larger context, more agent attempts, or the approach sometimes called tokenmaxxing or tokenmaxing needs a measure that connects consumption to a result an operating owner can judge.

Choose one bounded unit: a service-desk response accepted by its reviewer, a change that passes its agreed tests and review, or a case resolved under a defined policy. Record the acceptance rubric, task mix, observation window, and owner before comparing implementations. Count each distinct task at most once in the denominator, even if several attempts contributed to it. Keep failures, retries, cancellations, and rejected drafts in the cost numerator.

Anthropic's evaluation guidance distinguishes a task from a trial and from the resulting environment state. That distinction is useful here: a plausible answer or a successful tool call can be an intermediate step. Your acceptance criterion must describe the completed work.

The FinOps Foundation's unit-economics framework separates resource efficiency from business units. Cost per token helps investigate a system. Cost per accepted result helps decide whether its operating tradeoffs are worthwhile. The rubric still needs to protect quality, safety, and service levels.

02 · Cost scope

Count the whole run, and state what the number includes.

Cost per accepted outcome = scoped operating cost across all attempts ÷ distinct accepted tasks

For a practical operating comparison, include model charges, paid tools and runtime, and human review and rework. Convert recorded review minutes using a stated labor-rate assumption. Add shared operating costs only with an explicit allocation rule, and avoid counting a cost twice when a platform bill already includes it. The scorecard keeps these components separate so a buyer can inspect the scope.

A list-price token estimate is a planning input. The amount billed may reflect a subscription, negotiated rates, credits, batch processing, or other contract terms. Reconcile invoices for the same workload and period when claiming a billing change. Keep estimates, invoiced charges, and allocated labor assumptions visibly distinct.

Record missing coverage as missing. If tool charges or review time are unavailable, report a partial model-cost measure and name the gap; do not silently turn it into a fully loaded operating cost. With zero accepted tasks, cost per accepted outcome is undefined. Report the spend and failed-task count rather than dividing by one or presenting a zero.

Implementation investment and fixed costs belong in a separate decision view, with an agreed useful life and allocation basis if amortized. A lower unit operating cost alone proves neither a positive return on investment nor a cash saving. Finance needs evidence that the relevant expense changed or capacity was put to an agreed use.

03 · Worked scorecard

More model spend can still produce a cheaper accepted result.

Consider two hypothetical implementations for the same 100 service-desk tasks, assessed against the same rubric. Every figure below is synthetic arithmetic, not a customer result or a measured experiment. The assumed review rate is $60 per hour, and this operating view excludes fixed implementation costs.

Synthetic example · USD · same task cohort and acceptance rubric
MeasureBaselineCandidate
Distinct tasks100100
Attempts, including retries120150
Accepted tasks8090
Model cost$30$45
Tool and runtime cost$10$15
Human review and rework120 min90 min
Review cost at assumed $60/hour$120$90
Scoped operating cost$160$150
Cost per accepted outcome$2.00$1.67

The baseline is ($30 + $10 + 120 ÷ 60 × $60) ÷ 80 = $2.00. The candidate is ($45 + $15 + 90 ÷ 60 × $60) ÷ 90 = $1.6667, rounded to $1.67. Its assumed reduction in review effort outweighs its higher model and tool cost in this example. If review time had stayed at 120 minutes, the candidate would cost $180 ÷ 90 = $2.00 per accepted result.

The table illustrates sensitivity to the cost boundary and denominator. It does not establish that extra tokens caused better outcomes. In a real comparison, hold the task cohort and rubric stable, record configuration versions, inspect failed cases, and repeat on a representative held-out set. Report acceptance rate, latency, and safety exceptions alongside cost so a cheap failure cannot look like an improvement.

04 · Cache accounting

A cache-token share needs an explicit denominator.

Prompt caching reuses eligible input processing. It does not mean a prior answer was returned without new work. Anthropic and OpenAI expose different usage fields and caching behavior. Normalize each provider's fields into disjoint classes before summing: uncached input, cache creation where separately reported, cache reads, and output. When a provider reports total input inclusive of cache reads and cache creation, subtract both separately reported subsets to obtain uncached input; when input already excludes them, keep that field unchanged.

A separate synthetic token example has 900,000 cache-read tokens, 50,000 uncached input tokens, 25,000 cache-creation tokens, and 25,000 output tokens. Its cache-read share is 900,000 ÷ 1,000,000 = 90% of all counted tokens. Among input classes alone it is 900,000 ÷ 975,000 = 92.3%. Both numbers describe the same mix with different denominators.

Neither number is a request hit rate. If you define a hit as a request with any cached input, calculate requests with cached input ÷ eligible observed requests and state that definition. A request can have both cached and uncached input, and large requests can dominate token share. Token totals alone cannot recover how many requests hit.

A high cache-read share also does not establish the share of spend saved. Compute a rate-based estimate by summing each disjoint token class multiplied by its applicable rate, with the rate source, date, currency, provider, model, and terms recorded. Compare that estimate with a defined uncached counterfactual; reconcile actual invoices separately. Measure request and end-to-end task latency from timestamps. Token proportions cannot prove that a workflow became faster.

05 · Evidence

Join utilization to the task and its operating owner.

Agent Console, our MIT-licensed open-source tool, reads local Claude Code and Codex session evidence and presents token classes and list-price cost estimates. It is useful for investigating recorded coding-agent utilization. Its totals do not supply your task-acceptance rubric, human review time, or authoritative invoice on their own.

Agent Console Enterprise (ACE) connects evidence from supported provider, cloud, gateway, and coding-agent sources to ownership and business cases. Its public overview separates forecasts and provisional value from finance-confirmed evidence. It reads alongside the model request path; coverage depends on the sources connected. Treat both tools as evidence inputs to the scorecard, with their scope recorded.

For each task, retain an identifier, implementation version, attempt count, acceptance decision, and evidence location. Join usage and cost to that task where the source supports the relationship. Record unattributed spend and coverage gaps separately. A team total divided by a selected subset of successful tasks is misleading unless that allocation is explicitly justified.

06 · Use the worksheet

Make one decision with a reproducible comparison.

  1. Download the CSV and replace the synthetic values, scope, rubric, and evidence fields. Keep the formula rows; recalculate them in your spreadsheet. Missing, nonnumeric, or invalid required inputs return “incomplete”; a valid zero denominator returns “undefined”. The file uses a dot decimal and comma separators.
  2. Agree on the cost basis and coverage with engineering, the workflow owner, and finance. The sample is complete only within its stated synthetic operating scope.
  3. Compare like-for-like tasks. Record review time, retries, acceptance, latency, and safety exceptions for both configurations; leave missing measurements explicit.
  4. Review the evidence before expanding a workload. Revisit the scorecard when the model, prompt, tools, task mix, pricing terms, or acceptance rubric changes.

Our AI outcomes briefing develops the wider operating contract: who accepts the work, which business result matters, and who owns the system after release. This scorecard is the narrower accounting exercise that makes one implementation comparison inspectable.

Primary references

Reviewed 3 October 2026. The worksheet and worked arithmetic are LockedIn Labs' illustrative method; the sources support the accounting and evaluation distinctions.