AI Token Costs Are Making Your AWS Commitments Harder to Manage

Dr James Mitchell, Executive Leader, Strategic Blue
6th August 2026

Two years ago, only 31% of FinOps teams had a mandate to manage AI spend. Today it is 98% of the 1,192 respondents surveyed for the FinOps Foundation’s State of FinOps 2026 report, a sample representing more than $83bn of annual cloud spend.

That shift was one of the defining themes at FinOps X 2026, where the phrase that stuck was “tokens all the way down.” The token has become the atomic unit that almost every AI cost conversation eventually traces back to.

It is easy to read this as the arrival of an entirely new FinOps discipline. It is. But for most organizations, it is also a story about who inherits the work.

AI token cost management lands on the same FinOps and platform functions already responsible for AWS commitments, Reserved Instances, Savings Plans, coverage, utilization, forecasting, and executive reporting. Almost nobody gets a new team and a blank calendar for it.

Those teams stay lean. The same FinOps Foundation survey found organizations spending $100m or more a year on cloud average eight to ten practitioners plus three to ten contractors.

The 2026 State of AI in FinOps report, based on a Sapio Research survey of 700 engineering leaders and practitioners across the United States, United Kingdom, France, Germany and India in May and June 2026, found 67% of organizations now spend more than $250,000 per month on AI. One in five already spends more than $1 million each month. Respondents estimate 26% of all AI spend is wasted, which at $1 million a month is roughly $260,000 with no measurable return.

At that level AI has become a financial investment on the scale of core infrastructure. It needs the same governance, attribution, and forecasting that cloud computing took most of a decade to develop.

Organizations cannot afford to take another decade to build it.

The overlap between AI cost management and AWS commitment management therefore deserves closer attention, because it changes where limited internal FinOps capacity is best spent.

What makes AI token costs different from traditional cloud costs? 

AI token costs differ from traditional cloud costs in three ways: the unit economics, the ownership model, and the absence of a settled optimization playbook. Each one makes token spend harder to forecast and harder to attribute than compute hours or storage. That is why AI cost management deserves to be treated as its own discipline rather than a repackaged version of cloud cost governance.

The underlying unit economics are different.

Cloud costs are generally built around relatively familiar measures such as compute hours, storage consumption, and data transfer. AI costs are increasingly driven by tokens, model choice, request volumes, prompt length, output length, caching, and the architecture surrounding each application, whether it runs through Amazon Bedrock, a third-party model API, or self-hosted inference.

A single prompt-template change or a longer context window can move token consumption by an order of magnitude without touching the underlying infrastructure.

Ownership is also far less settled.

Cloud cost governance has had since 2009, when AWS introduced EC2 Reserved Instances, to build tagging standards, chargeback models, budgeting processes, and forecasting methods. AI spend has no comparable operating model yet.

The State of AI in FinOps report found that 52% of organizations have no clear owner for AI costs. Accountability is often split across engineering, platform teams, FinOps, finance, and IT.

Each function owns part of the problem, but nobody necessarily owns the complete bill.

Responsibility is also decentralized across the product and engineering teams building AI features. The decisions that drive spending can include:

  • Which model is selected
  • How prompts are structured
  • How much context is passed into each request
  • Whether responses are cached
  • How retrieval-augmented generation is designed
  • Which workloads run through managed APIs versus internal infrastructure
  • How quickly AI functionality is rolled out to users

There is no mature optimization playbook comparable to the one that exists for Reserved Instances and Savings Plans.

Model selection, prompt efficiency, context management, caching, and workload routing are all still evolving rapidly.

None of this reflects poorly on the teams managing AI spend. It is a genuinely new challenge.

The organizations handling it best share one habit. They put visibility, attribution, and forecasting in place before production usage scales, rather than waiting for the first unexpected invoice.

The sequence matters:

Visibility comes before attribution. Attribution comes before optimization.

Jumping immediately to “How do we reduce the bill?” without first understanding where the money is going and why rarely creates sustainable savings.

The same is true of forecasting.

Organizations that forecast AI spend using business context, such as planned model migrations, product launches, user growth, and new deployments, are much better positioned than those simply extrapolating from last month’s invoice.

There is a meaningful difference between:

“Our AI bill increased unexpectedly.”

And:

“Our AI bill increased because the RAG deployment planned for Q3 has entered production.”

The first creates concern. The second demonstrates control.

Why AI cost management lands on the same small FinOps team

AI cost management arrives on top of an existing workload rather than into spare capacity. The FinOps team that inherits token governance is the team already carrying the AWS commitment portfolio.

It is added to a list that already includes:

  • Cloud budgets
  • Forecasting
  • Tagging compliance
  • Showback and chargeback
  • Reserved Instance and Savings Plan coverage
  • Commitment utilization
  • Renewal timing
  • Cost anomaly investigation
  • Executive reporting

The State of AI in FinOps data reinforces the pressure these teams are already under.

More than half of respondents said AI budgeting remains largely guesswork, while 47% said they are constantly worried about their next token bill.

All of that lands on the same desks that already handle cloud cost management.

Adding AI visibility, attribution, forecasting, and optimization does not make anything already on the list easier. The team does not automatically grow because its responsibilities have expanded.

Headcount planning moves much more slowly than AI adoption.

The practical reality is that the same small number of people must absorb a fast-moving new discipline while continuing to manage everything they were already responsible for.

Renewals still come due.

Coverage still has to be checked against actual usage.

Utilization still drifts when nobody is watching.

Management still expects its report.

This raises a question organizations should answer deliberately, rather than allowing it to resolve itself through neglect:

If FinOps capacity is fixed but the scope of responsibility has doubled, which discipline requires the greatest level of internal attention, and which could be managed more effectively by a specialist partner?

Where AI costs and AWS commitments overlap

It would be convenient if AI token management and AWS commitment management stayed in separate lanes.

They do not. The two converge the moment training or inference workloads move onto reserved GPU and Trainium capacity.

AI workloads introduce new complexity directly into the AWS commitment environment.

Training clusters, inference endpoints, and supporting infrastructure bring new instance families, new usage patterns, and new commitment decisions into a portfolio that the FinOps team was already managing.

AI infrastructure becomes part of the AWS commitment portfolio, not a separate one beside it.

Training and fine-tuning workloads increasingly rely on GPU and Trainium-accelerated infrastructure, including P5 (NVIDIA H100), P6 (NVIDIA Blackwell), and Trn-series capacity.

That capacity can be reserved through mechanisms designed specifically for AI workloads, such as:

  • EC2 Capacity Blocks for ML
  • SageMaker training plans
  • GPU and Trainium Reserved Instances
  • Compute Savings Plans

These instruments behave differently from traditional commitment models in several important ways.

Shorter, scheduled usage windows

Traditional commitments are often based on steady-state workloads that remain relatively consistent over months or years.

AI training workloads can be much more variable.

Capacity Blocks are designed for scheduled training runs and can be reserved in advance for specific windows, often lasting weeks or months rather than multiple years.

This changes the nature of the forecasting problem. Teams must predict not only how much capacity they need, but precisely when they will need it.

Faster hardware cycles

A reservation tied to a particular GPU or Trainium generation can remain linked to that generation for the entire commitment period.

That matters because AI infrastructure is evolving quickly.

New chips, instance families, and model architectures are appearing faster than the traditional three-year commitment cycle was designed to accommodate.

A commitment that looks appropriate today can become restrictive as newer hardware becomes available.

Price movement after commitment

AI infrastructure pricing can also change significantly during the commitment term.

On 5 June 2025, AWS cut pricing on its NVIDIA GPU instances: up to 45% on P5, up to 33% on P4d and P4de, and up to 26% on P5en. The reduction applied to On-Demand from 1 June and to Savings Plan purchases made after 4 June. A week later AWS cut Amazon SageMaker AI P4 and P5 pricing by up to 45% as well.

Teams that had already committed at previous price levels could therefore find themselves holding an effective rate above the newer market price.

Commitments are still worth making. Flexibility and timing simply carry more weight than they used to.

Compute Savings Plans as a bridge

Compute Savings Plans can apply across standard EC2, GPU, and Trainium families.

This makes them an important bridge between traditional commitment management and AI-era infrastructure.

They provide greater flexibility than single-instance commitments and can reduce some of the risk created by faster hardware cycles and changing workload requirements.

However, the discipline of managing coverage and utilization across AI infrastructure is still newer and less established than it is for conventional compute.

Multi-payer complexity

AI workloads also frequently land in new AWS accounts.

Those accounts may be separated for security, attribution, governance, or organizational reasons. In larger businesses, these accounts are often spread across multiple payer environments, so building the full commitment picture means aggregating data across every payer.

This fragments the commitment picture further.

AWS-native tools such as Cost Explorer and Cost Anomaly Detection operate within individual payer environments. They do not provide a consolidated view across every payer account an organization operates.

As a result, the FinOps team may need to manually piece together:

  • Commitment coverage
  • Utilization
  • Renewal dates
  • Remaining liability
  • Exposure across accounts
  • Changes in instance mix

The same capacity constraint is therefore being pulled in two directions.

Token visibility has become a new, urgent, and increasingly board-visible priority. At the same time, the AWS commitment environment is becoming more complex rather than remaining static.

Reserved GPU capacity solves a genuine business problem. Access to high-demand hardware can determine whether a training run happens on schedule.

AI infrastructure commitments are worth making.

What they do not do is reduce the work of managing AWS commitments. They open a second, faster-moving front on top of the Reserved Instance and Savings Plan environment that already existed.

AI-era commitments vs. traditional EC2 commitments

  Traditional EC2 commitments AI and GPU-accelerated commitments
Typical usage pattern Steady-state and predictable over months or years Often variable, with scheduled training runs and changing inference demand
Common mechanisms Reserved Instances and Savings Plans EC2 Capacity Blocks for ML, SageMaker training plans, GPU or Trainium Reserved Instances, and Compute Savings Plans
Typical term One or three years Capacity Blocks may cover weeks or months, while RIs and Savings Plans generally remain one or three years
Primary risk Usage falling below the committed level Hardware-generation lock-in, price changes, and rapidly shifting workload requirements
Flexibility lever Convertible RIs and Compute Savings Plans Compute Savings Plans, shorter reservations, and avoiding long single-instance commitments
Visibility challenge Fragmented across accounts and payer environments Further fragmentation as AI workloads move into dedicated accounts
Typical owner FinOps or platform team Often the same team, now also responsible for AI infrastructure and token attribution

What happens when AWS commitment management gets deprioritized

When AWS commitment management slips down the priority list, coverage and utilization drift and existing savings erode. That is the predictable cost of AI cost visibility becoming the urgent problem while FinOps capacity stays fixed.

The shift in attention is understandable. A genuinely new and board-visible challenge will always pull harder than a portfolio that looks stable.

There is usually no pricing event behind the erosion.

The commitment strategy is usually sound as well.

Monitoring simply becomes less frequent.

The State of AI in FinOps report found that 79% of organizations need at least a full day to trace a cost spike to its source. Nearly one-third need an entire week.

For an organization spending $1 million a month, a week of investigation covers about a quarter of a monthly bill, roughly $230,000, before anyone knows what moved.

When that investigative workload competes with existing FinOps responsibilities, commitment management can become the area that gets deprioritized.

This rarely happens as an explicit decision.

It happens gradually.

Renewals are assessed using old assumptions. Usage patterns shift without the commitment portfolio being adjusted. Coverage declines. Utilization drops. Savings erode.

The risk is greater today because the AWS environment itself is more complex.

A commitment portfolio containing GPU, Trainium, and inference capacity needs more active oversight, not less, at exactly the moment internal attention is being pulled toward token governance.

How managed rate optimization keeps the AWS commitment engine running

Managed rate optimization moves the repeatable part of AWS commitment management to an external specialist, which frees internal FinOps time for AI cost governance.

Strategic Blue’s role is to keep the AWS commitment engine running, including:

  • Coverage
  • Utilization
  • Renewal timing
  • Instrument selection
  • Commitment flexibility
  • Cross-payer visibility
  • Changes in workload and instance mix

This includes the additional complexity introduced by AI-driven EC2, GPU, and Trainium usage.

The objective is straightforward: maintain and improve commitment performance while requiring minimal ongoing time from the customer’s FinOps and engineering teams.

In practice, Strategic Blue uses read-only access to cost and usage data through a CloudFormation stack.

Setup takes approximately 15 minutes and requires:

  • No changes to infrastructure
  • No tagging project
  • No new platform for engineering teams to learn
  • No resale of the existing AWS relationship

Strategic Blue then manages commitment coverage across the customer’s payer environments and consolidates the environment into a single view.

None of this makes AI cost management simple.

It is a decision about where internal attention creates the most value.

AI cost governance is still evolving. It benefits from close collaboration between product, engineering, finance, and FinOps. Internal teams need to develop new practices around token attribution, model selection, prompt efficiency, deployment architecture, and forecasting.

AWS commitment management is more mature.

The questions are better understood:

  • When should a commitment be renewed?
  • How much coverage is appropriate?
  • Which instrument offers the right level of flexibility?
  • How should the portfolio respond when the instance mix changes?
  • Where is commitment liability building?
  • How should multiple payer environments be managed together?

A specialist team answers these questions every working day, across customer estates of different sizes, instrument mixes and payer structures.

That creates a level of pattern recognition and market awareness that is difficult for an individual organization to reproduce internally.

Why not simply hire another FinOps engineer?

Hiring another FinOps engineer looks like the obvious answer to a doubled scope. It usually is not.

AI cost management and AWS commitment management are moving in increasingly different directions.

AI cost governance requires close proximity to engineering and product decisions. Model selection, prompts, caching, deployment architecture, and user behavior can change weekly.

Commitment management requires detailed knowledge of AWS pricing instruments, portfolio-level analysis, renewal discipline, and continuous monitoring across accounts.

A single generalist covering both areas risks becoming reactive in both.

A more effective model is often:

Dedicated internal attention on AI cost governance, supported by specialist external management of AWS commitments.

This gives each discipline the focus it needs.

Strategic Blue has more than 14 years of cloud cost forecasting experience and is both an AWS Advanced Tier Services Partner and a FinOps Foundation-certified Specialty Solution Provider.

That experience is increasingly important as AI infrastructure adds faster hardware cycles, new commitment mechanisms, and greater volatility to the AWS environment.

Why use an independent commitment partner?

An independent commitment partner has no revenue interest in how much you commit, which is the structural reason not to rely entirely on AWS-native recommendations.

AWS benefits when customers increase their committed spend. Greater commitments create more guaranteed platform revenue.

An independent advisor is instead focused on optimizing the shape of the commitment portfolio:

  • The appropriate commitment level
  • The right term length
  • The balance between savings and flexibility
  • The most suitable instrument
  • The ability to respond as workloads change

Sometimes that means committing more aggressively.

Often it means committing less with the easiest commitment instrument, and adding a different type of flexibility by adding a second instrument.

This blend of commitment instruments preserves options while AI infrastructure and pricing keep moving.

Incentives shape recommendations, and a platform’s do not point the same way as a customer’s.

The goal is maximum sustainable savings at a controlled level of risk, which is a different target from maximum commitment.

Let your FinOps team focus on the new problem

AI cost management is quickly becoming one of the most important responsibilities within FinOps.

It is also one of the least mature.

Teams need time to establish visibility, ownership, attribution, and forecasting before they can optimize effectively.

That work requires close internal attention.

AWS commitment management still matters just as much as it did before. In many organizations, the introduction of AI infrastructure is making it more complex.

Both disciplines need focussed attention without frequent context-switching between the two.

The decision worth making deliberately is where the internal capacity goes.

By handing mature, repeatable commitment management to a specialist partner, FinOps teams can focus their attention on the new challenge in front of them without allowing the savings already built into their AWS environment to quietly erode.

Frequently asked questions

Is AI cost management the same as AWS rate optimization?

No. AI cost management and AWS rate optimization are separate disciplines.

Token costs and cloud compute costs have different unit economics, ownership structures, and optimization methods.

However, they share a common constraint: the time and capacity of the FinOps team.

AI workloads also introduce new GPU and Trainium commitments into the AWS environment, meaning the two disciplines are increasingly connected.

Who typically owns AI token costs?

Ownership is still developing.

The 2026 State of AI in FinOps report found that 52% of organizations have no clearly defined AI cost owner.

Responsibility is often divided across engineering, platform teams, FinOps, finance, and IT.

In practice, it frequently defaults to the team already responsible for cloud cost governance, although product and engineering teams are often closest to the decisions driving usage.

What happens if nobody actively manages Reserved Instances and Savings Plans?

Coverage and utilization rarely collapse immediately. They drift.

Usage patterns change, commitments renew against outdated assumptions, and savings gradually begin to erode.

The earlier this drift is identified, the easier and less expensive it is to correct.

Are GPU and Trainium commitments riskier than standard EC2 commitments?

They carry different risks.

GPU and Trainium reservations can be tied to a particular hardware generation, while new instance generations and price changes may arrive faster than the commitment term.

Compute Savings Plans can reduce some of this risk by providing coverage across different compute families, including GPU and Trainium infrastructure.

What is the difference between EC2 Capacity Blocks for ML and SageMaker training plans?

Both allow organizations to reserve accelerated computing capacity in advance, but they operate at different layers.

EC2 Capacity Blocks reserve raw EC2 capacity directly.

SageMaker training plans reserve capacity within the managed SageMaker environment and can support training jobs, HyperPod clusters, and certain inference workloads.

The right option depends on whether the workload runs directly on EC2 or through SageMaker.

Does reserving AI infrastructure guarantee lower costs?

No. Reserving AI infrastructure does not guarantee a lower bill.

Reserved GPU and Trainium capacity can solve an availability problem as much as a cost problem.

It secures access to high-demand infrastructure for a defined period.

Whether it also reduces costs depends on how accurately the commitment matches actual usage throughout the term.

Why not hire one FinOps engineer to manage both areas?

The two disciplines require different kinds of expertise.

AI cost management depends heavily on engineering and product decisions such as model selection, prompts, caching, and deployment design.

Commitment management requires detailed knowledge of AWS pricing instruments, coverage, utilization, renewals, and portfolio risk.

One individual covering both areas may struggle to develop sufficient depth in either.

How does managed rate optimization work without engineering effort?

A CloudFormation stack provides Strategic Blue with read-only access to the relevant AWS cost and usage data.

There are no infrastructure changes, no tagging requirements, and no new tools for engineers to adopt.

Strategic Blue then manages commitment decisions, including instrument selection, coverage, term length, renewal timing, and ongoing monitoring.

Sources

  • FinOps Foundation, State of FinOps 2026 Report
  • FinOps Foundation, FinOps X 2026
  • Harness, 2026 State of AI in FinOps Report
  • AWS, Amazon EC2 Capacity Blocks for ML
  • AWS Documentation, Reserve Capacity With SageMaker Training Plans
  • AWS, Pricing and Usage Model Updates for EC2 Instances Accelerated by NVIDIA GPUs
  • AWS, Announcing Price Reductions for Amazon SageMaker AI GPU-Accelerated Instances
[gtm]