Token costs in LLM agent pipelines are not a billing artefact to be reconciled at month-end. They are a runtime variable that compounds across every step of execution, and teams that treat them as the former will consistently discover that their cost models are wrong by the time they matter. The core problem is architectural: most agent pipelines are built to optimise for task completion, not for cost predictability, which means the economics of a given workload are only legible in retrospect. This article explains why that happens, what the structural causes are, and how engineering teams can build cost forecasting directly into the execution path rather than bolting it on afterwards.
Companion piece to our broader work on agent cost governance. See AI Coding Agent Costs: Token Budget Governance Guide for how to govern token spend organisationally before it spirals.
Why Token Consumption Varies by an Order of Magnitude Across Identical Tasks
The variance is not noise. It is a structural property of how agents execute. An agent resolving a software engineering task may converge in a single edit or may require multiple rounds of tool calls, error handling, and re-planning before it reaches a valid state. Each intermediate step appends output to the shared context, which is then re-read in full by every subsequent call.
This compounding effect is the mechanism behind the observed variance. A task that takes three steps instead of one does not consume three times the tokens. It consumes substantially more, because the growing context inflates the input size of every later call in the sequence. Research on this problem confirms that token consumption can vary by over an order of magnitude across runs of the same task (Ouyang et al., arXiv 2026).
The commercial implication is direct. If your cost model assumes a fixed token budget per task type, it will be accurate on average and wrong on the runs that matter most, specifically the long-tail runs that drive the majority of actual spend. A budget set to the median will be exceeded regularly. A budget set to the 95th percentile will be wasteful on most runs and may still be breached on outliers.
The Composable Cost Representation Approach
The key insight from recent work on this problem is that predicting total token consumption requires modelling not just what a given step consumes, but what context growth it introduces for all subsequent steps. A step that produces verbose tool output is expensive not only in itself but in the compounding input cost it creates downstream.
TokenCast, proposed by Ouyang et al. (arXiv 2026), formalises this as a composable cost representation. Each execution segment is characterised by two quantities: its own token consumption and the context growth it contributes. Composing adjacent segments produces a cumulative estimate that accounts for the re-reading cost incurred when earlier context is carried forward into later calls.
This framing is useful for engineering teams because it maps directly onto the structure of an agent trace. Rather than treating the run as a single opaque unit, it decomposes cost into segment-level primitives that can be estimated, updated, and composed incrementally as execution proceeds.
Building Runtime Estimation Into the Execution Pipeline
Segment-Level Estimation Without Additional LLM Calls
The practical constraint on any runtime cost estimator is that it cannot itself be expensive to run. An estimator that requires additional LLM calls to produce a forecast defeats its own purpose. The TokenCast approach addresses this by using lightweight learned models over execution segment features, achieving a mean cumulative prediction time of 32.8 milliseconds per run on SWE-bench Verified (Ouyang et al., arXiv 2026). That is within the latency budget of a synchronous execution loop.
For engineering teams, this means the estimator can sit inline in the agent execution loop, updating its forecast after each step completes, without adding meaningful overhead to wall-clock latency.
Evidence Refresh as Execution Unfolds
Static pre-execution estimates are useful for budget allocation but insufficient for runtime control. The estimate that matters for a given run is the one conditioned on what has already happened. As each step completes, the observed segment cost and context growth become evidence that should update the forecast for the remaining steps.
This evidence-refresh pattern is what makes runtime cost estimation operationally useful. It allows the pipeline to detect early in a run whether the trajectory is trending toward a high-cost outcome, rather than discovering that only at the end. Teams can then implement intervention logic, such as context truncation, early stopping, or task rerouting, at a point where it still affects the outcome.
Operational Implications for Budget Control
Fixed Budgets Versus Adaptive Budgets
A fixed per-task token budget is the simplest control mechanism, but it is also the bluntest. Applied uniformly, it will terminate runs that were close to completion and waste budget on runs that could have been resolved cheaply. In offline replay experiments, a dynamic budget policy informed by TokenCast forecasts used 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion rates (Ouyang et al., arXiv 2026). That is a material efficiency gain achievable without changing the underlying agent or model.
The mechanism is straightforward. An adaptive budget policy can allocate more tokens to runs that are making measurable progress and intervene earlier on runs that are expanding context without converging. This requires a runtime signal, which is exactly what a segment-level cost estimator provides.
Integration With Model Routing Economics
Cost forecasting also interacts with model routing decisions in ways that are easy to underestimate. Switching to a cheaper model mid-run can reduce per-token cost, but if the switch happens at a context boundary where the new model must re-read a large accumulated context, the savings may be smaller than expected. We have written separately about how intelligent model switching can produce counterintuitive cost outcomes in agent pipelines. Cost forecasting that accounts for context growth makes these routing decisions legible before they are made, not after.
What Engineering Teams Should Build First
The entry point is not a full forecasting system. It is instrumentation. Teams that do not have per-step token consumption logged at the segment level cannot build cost models, because they have no training signal. The first investment is ensuring that every agent execution emits structured logs that capture input tokens, output tokens, and context size at each step boundary.
From that foundation, a lightweight segment-level estimator can be trained on historical traces. The model does not need to be complex. The goal is to produce a forecast that is accurate enough to support budget control decisions, not to predict the exact token count. Even a rough estimate that correctly identifies the top decile of cost runs before they complete has significant operational value.
The broader organisational shift is treating token forecasts as a first-class output of the execution pipeline, alongside latency and error rate. Teams that instrument this way gain the ability to set meaningful per-task budgets, detect cost anomalies in real time, and make routing and truncation decisions on the basis of evidence rather than heuristics.
Where Vector Labs Fits
We build production agent systems with cost governance designed into the execution architecture, not retrofitted after deployment. In our token governance analysis, we set out the organisational and architectural patterns that prevent agent token spend from compounding into unmanageable infrastructure debt. If your team is running agents at scale and needs a structured approach to cost predictability, contact us at vector-labs.ai/contacts.
FAQs
The variance comes from two compounding effects. First, the number of steps an agent takes is determined by tool feedback and intermediate results, which differ across runs even for identical starting conditions. Second, context accumulates across steps, so each additional step inflates the input size of every subsequent call. A task that takes twice as many steps does not cost twice as much in tokens - it costs substantially more because of this compounding input inflation.
Both are useful for different purposes. Pre-execution estimates, based on task type and historical segment distributions, support budget allocation and capacity planning. Runtime estimates, updated after each step using observed segment costs and context growth, are what enable active budget control decisions during execution. The most effective approach uses pre-execution estimates to set initial budgets and runtime estimates to adjust those budgets as evidence accumulates.
Not if the estimator is built correctly. Lightweight models operating over structured segment features can produce updated forecasts in tens of milliseconds per step, which is well within the latency budget of a typical agent execution loop. The key constraint is that the estimator must not require additional LLM calls - those would add both latency and token cost, undermining the purpose of the system.
At minimum, every agent execution needs to emit structured logs capturing input token count, output token count, and cumulative context size at each step boundary. Without per-step data, you can only model total run cost, which is too coarse to support segment-level forecasting or early intervention. Most teams find that adding this instrumentation also surfaces unexpected patterns in how context grows across different task types, which is useful independent of cost modelling.
Model routing and cost forecasting are closely coupled in practice. Switching to a cheaper model mid-run can reduce per-token cost, but if the switch occurs at a point where the accumulated context is large, the cheaper model must still read that context in full. A cost estimator that accounts for context growth allows routing decisions to be evaluated against the full cost trajectory rather than just the per-token rate of the candidate model, which produces more accurate routing economics.

