Strategy
AI FinOps: why agent costs escape budgets, and how to control them
Token prices are falling, but an agent chains ten to twenty calls per task. Cost per task, not price per token, becomes the indicator a CTO must manage.
A paradox recurs in 2026 AI cost studies: token prices fall year over year, yet bills keep rising. According to several industry syntheses, a large majority of companies overshoot their AI cost forecasts, and nearly all FinOps teams are now handed the topic. These figures come from surveys and compilations of uneven quality; keep the order of magnitude and direction, not the decimal.
Why a falling unit price isn't enough
A chatbot request is one call. An agent task is a loop: planning, tool call, reading the result, retrying on failure, verifying. Industry analyses cite ten to twenty calls per task instead of one, with contexts growing longer at each step. Cost per task then depends on what the model decides to do, and can vary by an order of magnitude between two executions that look identical. Classic budget forecasting — users × average cost — stops working.
Five architecture levers
- Route by difficulty. Send simple tasks to small, cheap models, and reserve powerful models for the steps that warrant them. A well-tuned router changes system economics more than negotiating a rate.
- Cache. Stable system prompts and contexts, frequent answers: prefix and result caching reduces both cost and latency.
- Cap loops. Maximum steps, a token budget per task, stop with human escalation beyond it. An agent without a cap is a bill without a cap.
- Shrink context. Don't resend the full history at every step: summarize, retrieve only what's needed. It also improves answer quality.
- Evaluate self-hosting or open models for stable, predictable volumes, with the hidden costs: operations, GPUs, security, skills. The sovereignty argument may align with the cost argument, but compute both.
Measure the right thing
The useful indicator isn't total monthly cost but cost per successful task, set against the task's value. It requires three things: attributing each call to an agent, a use case and a team (tagging and logging from the design stage); measuring success rate, since a failed task costs as much as a successful one; and comparing with the alternative, human or classic software. Some use cases won't pass this test — and that's information the CTO needs before production, not after.
The link with governance
Gartner names escalating costs among the three main causes of agentic project cancellation (see our article on agents in production). AI FinOps is therefore not a peripheral finance topic: it is an architecture requirement, like observability. It is set up with it — same traces, same tags.
- Token price falls, but agents multiply calls: total cost can rise even as unit price drops.
- Difficulty routing, caching, loop caps, smaller context and open models are the architecture levers.
- The indicator to manage is cost per successful task, relative to its value.
- Tagging and traces are designed in from the start: it is the same infrastructure as observability.
Sources
Related reading
Does this challenge sound familiar?
A first conversation to assess it together, at no cost.