Agentic AI · Cost & ROI

Token prices fell.
The bill didn't.

An advanced model on the cusp of innovation saw a 67% decrease in cost annually. Despite this, Enterprise AI expenses continued to increase due to the high token consumption of agentic workflows, which require five to thirty times more tokens per task compared to a simple chat reply. The total expenditure is determined by both price and volume, not just price alone, ultimately shaping the economics of agents in 2026.

overview SLIDE 01 : AGENT ECONOMICS Cover slide: Agent Economics
Quick answer

Agent economics Understanding, measuring, and managing the costs and value of running AI agents is a new challenge brought about by agentic AI, which has transformed enterprise AI spending from a fixed, predictable expense to a variable, consumption-driven compute cost. Tokens have become the atomic unit of that economyAs every plan, tool, and decision burns them, total spending continues to rise despite falling per-token prices because agentic workflows require more tokens per task than a basic chat reply. This shift has moved the focus from cost-per-token optimization to measuring. cost per verified outcomeAccording to McKinsey, well-planned enterprise deployments can achieve a 5.8x ROI within fourteen months, with success attributed to economically disciplined orchestration rather than the size of the model.

▲ $2/1M in▼ 67% YoY▲ 30x complexity▼ token price▲ 5.8x ROI
▲ cost shift SLIDE 02 : FIXED LABOR COST TO VARIABLE COMPUTE The shift from fixed software and labor costs to variable, consumption-driven agent compute costs
A 30x jump, in three years

Simple software costs; orchestrated agents consume

In 2023, a basic linear workflow cost around four cents per interaction. By 2026, a more intricate orchestrated system with additional tools, reasoning, and iterative loops costs approximately a dollar twenty per interaction - about thirty times higher for a system performing significantly more work, not just the same work at a slower pace.

Boards must understand the change: AI now trades a fixed software license or labor cost for a dynamic compute consumption model, more like a utility bill than a subscription. The total spend is determined by the price per token multiplied by the volume consumed, with volume being the variable often overlooked in budgets.

▼ price/token▲ volume▼ predictability▲ token demand
The atomic unit of value

Tokens are the currency; the exchange rate keeps moving

A token is a compact unit of processed data, such as text characters, image fragments, or audio snippets. Every action, search, and decision made by an agent is quantified in terms of tokens.

tokenomics SLIDE 03 : TOKENS AS THE UNIT OF AI ECONOMICS Tokens as the atomic unit of AI economics, and why falling prices don't mean falling bills
A cheaper token, not a cheaper bill

Even when unit price drops, aggregate spend climbs

The decline in per-token list prices is overshadowed by organizations increasing modality, agent autonomy, and reasoning chains, leading to a decrease in per-unit savings. After implementing multi-agent systems, a large enterprise saw their daily token usage increase from 8 billion to 27 billion, resulting in individual tokens becoming cheaper but overall token costs remaining the same.

▲ tiered routing▼ 87% gap▲ frontier premium
▲ savings lever SLIDE 04 : MODEL ROUTING & CASCADE ARCHITECTURE Model routing and cascade architecture: matching model tier to task complexity
Not every step needs the flagship model

Reserve frontier pricing for frontier work

Classifying, extracting, detecting intent, and summarizing documents are the standard processes in enterprise workflows that do not necessitate advanced capabilities. The cost difference between basic and advanced models ranges from 20 to 50 times per token. Routinely sending all tasks to the top-tier model, regardless of complexity, is a major cause of excessive spending.

In a large-scale analysis, organizations that implement a tiered architecture, with small models for routing and classification, a mid-tier for structured reasoning, and a frontier for high-value or high-risk decisions, had a median blended cost of $2.31 per million tokens. In contrast, organizations routing all tasks to frontier models incurred a cost of $18.40 per million tokens, resulting in an 87% cost difference. This highlights the significant impact of architectural decisions made early in deployment, which are often not revisited.

▼ cost/token▲ cost/outcome▼ vanity metric
the real KPI SLIDE 05 : COST PER OUTCOME, NOT COST PER TOKEN Cost per verified outcome as the real economic metric, replacing cost per token
Executives don't care about token efficiency

They care about what one result costs

Cost per resolved claim, cost per processed invoice, and cost per deployed feature are the key metrics that drive a business's operations. For an agent, unit economics can be simplified to total operating cost divided by completed work items that meet quality standards, rather than simply dividing raw token spend by other factors.

Companies that prioritize aligning metrics with outcomes are more successful than those solely focused on maximizing inference efficiency. A workflow may appear cost-effective per token, but could prove to be a poor investment if a low percentage of outputs meet review standards - the opposite is also true.

▲ $60M saved▲ $3.5B saved▲ 5.8x ROI▲ 192% avg
Measuring the return, honestly

Net value, not just time saved

Only considering saved time can inflate ROI and prioritize the wrong agents. A comprehensive formula calculates value against all actual costs incurred by the agent.

▲ net value SLIDE 06 : THE AGENT ROI FRAMEWORK Agent ROI framework: net value formula including costs, failures, and support
Value minus every real cost, not just the obvious one

Time saved, leakage prevented, revenue safeguarded - all efforts subtracted

The net value is determined by subtracting model cost, tool cost, integration cost, human review cost, failure cost, and support cost from the sum of time removed, leakage avoided, revenue protected, and quality gained. The elimination of minutes of genuine human work is a key factor in determining the overall net value.

5.8x
ROI within 14 months, well-scoped deployments
192%
Average ROI, US enterprises
74%
Of orgs see ROI within year one
$3.5B
Reported savings, one enterprise deployment
▼ 25-35% build▲ 65-75% run▼ upfront cost
TCO SLIDE 07 : TOTAL COST OF OWNERSHIP Total cost of ownership for AI agents: build cost is a fraction of the three-year total
The build is the smaller number

Build cost is 25–35% of the real three-year bill

The upfront cost of an agent's initial build usually makes up only 25-33% of the total three-year ownership cost. For instance, an $80,000 build suggests a three-year budget nearing $600,000 when factoring in LLM consumption, infrastructure, maintenance, monitoring, and human oversight.

The decision between building or buying follows a similar logic: custom infrastructure is typically more cost-effective and capable than a packaged SaaS model by the second year, but only becomes justifiable when token production reaches a sufficient scale.

▲ completion rate▼ iterations▲ resolution rate
What to actually watch

Traditional AI metrics don't capture enterprise economics

Autonomous systems require KPIs that accurately reflect how they allocate both time and money, rather than relying solely on accuracy metrics suited for a one-time model.

KPI panel SLIDE 08 : OPERATIONAL KPIs FOR AGENT ECONOMICS Operational KPIs for agent economics: completion rate, resolution rate, iterations, cost per result
Four numbers, watched together

High numbers of iterations indicate a cost issue, not a measure of quality.

A high number of reasoning iterations per workflow indicates inefficiency and excessive token usage, rather than thoroughness. The key is to create concise execution loops with strict iteration limits, instead of allowing the agent to endlessly reflect in the pursuit of quality.

Reliability
Completion rate

% of workflows completed correctly, end to end, without correction.

Labor displacement
Autonomous resolution rate

% resolved without any human intervention.

Cost signal
Reasoning iterations

High loop counts in workflows indicate inefficiency, not attention to detail.

The metric that matters
Cost per resolved outcome

What one completed, review-passing result actually costs.

▼ spend ceiling▲ chargeback▼ shadow AI
▼ cost controls SLIDE 09 : FINOPS PRACTICES FOR AGENTIC AI FinOps practices adapted for agentic AI: monitoring, chargeback, ROI thresholds, kill switches
The function that governed cloud now governs tokens

The bill comes after the damage is already done if there are no circuit breakers.

In just one year, the responsibility for overseeing AI expenses shifted drastically among FinOps professionals, with nearly all practitioners now managing these costs. This new challenge involves navigating a complex cost structure, characterized by token-based, consumption-driven, and constantly changing architecture, for which there is no established playbook.

Implementing practices such as real-time consumption monitoring at the workload level, business-unit chargeback for token use, ROI thresholds for new initiatives, explicit policies for discovering and remediating shadow AI, and hard kill switches at various levels is essential to prevent financial damage before it occurs.

▲ discipline wins SLIDE 10 : ECONOMIC DISCIPLINE OVER RAW CAPABILITY Summary: economic discipline in agent orchestration matters more than raw model capability
The competitive edge in 2026

The best model doesn't win. The best-run one does

In 2026, the company that will have a competitive edge is not the one with the most cutting-edge model, but the one with the most economically disciplined approach, including tiered routing, iterative loops, real-time consumption tracking, and evaluating deployments based on cost per verified outcome rather than cost per token.

Organizations that focus on measuring tokens will hit a plateau, while those that measure outcomes will experience growth. Begin with a single workflow, accurately calculate the full net-value formula, and base the decision to scale on the cost per verified outcome, rather than enthusiasm alone.

Frequently asked

Agent economics FAQ

Total spend is determined by multiplying the price per token by the volume consumed, with the volume growing faster than the price is decreasing. Agentic workflows require significantly more tokens per task compared to a basic chat interaction, and as enterprises increase the complexity and length of their reasoning chains, any potential savings per unit are overshadowed.

Using model routing, also known as cascade architecture, involves sending basic tasks such as classification and extraction to smaller, more cost-effective models, while saving more advanced models for important or risky decisions. A study discovered that organizations implementing tiered routing had an average cost of $2.31 per million tokens, compared to $18.40 for those relying solely on frontier models : an 87% cost difference resulting from a strategic architectural choice.

Executives and businesses prioritize cost per result, such as cost per resolved claim and cost per processed invoice, over the efficiency of token usage alone. A workflow may appear cost-effective per token, but could still be a bad investment if too few outputs meet quality standards; unit economics should be assessed as total operating cost divided by completed, review-passing work items.

The net value is the result of time saved from genuine removal, avoided leakage, protected revenue, and gained quality, minus various costs such as model, tool, integration, human review, failure, and support costs. Focusing solely on time saved without considering these costs can lead to an overestimation of ROI and the potential scaling of agents that may not be economically viable.

Usually, just 25-35% of the overall cost of ownership over three years is allocated towards the initial build. The rest is attributed to expenses such as LLM API usage, infrastructure, maintenance, monitoring, and human supervision. This is why simply budgeting for the build alone will significantly underestimate the true lifetime cost of running an agent.

Tracking usage in real-time at the workload level, charging business units for token use, implementing ROI thresholds for new initiatives, creating a policy for discovering shadow-AI, and implementing spend ceilings, call-volume caps, and automatic shutoffs at various levels.

Get started

Ready to model your agent's real economics?

Develop a tiered routing architecture prior to scaling up, consider the complete three-year total cost of ownership rather than just the initial build cost, and evaluate each deployment based on cost per verified outcome.