AI Cost Optimization on AWS: How to Measure What Actually Matters

Max
2nd October 2026

AI is getting cheaper. At least, that's what the headlines would have you believe.

The price of access to ever more capable models keeps falling, yet plenty of organizations building with generative AI are watching their total AI bill head the other way. Adoption is spreading, workloads are getting more sophisticated, and agentic systems now call models at machine speed. Keeping that spend in check has become one of the trickier problems in FinOps.

So how should you actually measure AI cost efficiency on AWS?

At a recent AWS session, Yuriy Prykhodko, Principal Technologist at AWS Enterprise Support and founder of the Cloud Intelligence Dashboards project, walked through a framework he developed alongside some of AWS's largest and most complex AI customers. The idea at its heart is simple: cheap tokens aren't the goal, efficient value is.

Here's how to apply that thinking to Amazon Bedrock and the rest of your generative AI estate.

Why cheaper AI models don't necessarily mean lower AI costs
It's tempting to assume that falling token prices will pull AI spend down with them. In practice they often don't, for two reasons.

1. Cheaper units can drive more consumption
Economists call this the Jevons paradox. When something becomes cheaper and more efficient to use, demand can rise by enough that total consumption goes up rather than down.

Generative AI fits the pattern unusually well. Traditional software tends to make an API call when a person does something. Agentic workflows don't wait for a person: an agent reasons, calls a tool, checks the result, and kicks off the next step, generating whole chains of model calls on its own.

In other words, LLMs are increasingly being consumed at the pace of machines rather than people. A lower cost per token can sit quite comfortably alongside a much bigger overall bill.

2. More capable models can use more tokens
Reasoning models add a second complication. A user might get a short, tidy answer while the model has done a great deal of processing behind the scenes to produce it.

So the useful question isn't really "How cheap is this model?" It's "How efficiently are we turning AI consumption into useful outcomes?" Answering that takes a more mature approach to generative AI FinOps.

Start with the foundations of AI cost management
Before you start building sophisticated AI cost KPIs, you need three basics in place.

1. AI cost allocation
First, you need to know who is consuming AI and which workloads are behind the spend. Depending on how your organization is set up, that might mean breaking Amazon Bedrock costs down by application, product, team, business unit, project, customer, environment, or workload.

Bedrock now offers several ways to attribute inference costs, including IAM principal attribution, application-level approaches, and request metadata. IAM principal identity can flow through to AWS Cost Explorer and CUR 2.0, which gives you much clearer sight of which users, roles, teams, or cost centers are driving spend.

Without allocation, a rising Bedrock bill is just a number. With it, you can start asking better questions. Which application is behind the increase? Which team is using the most AI? Which workloads are running on more expensive models? That's usually where optimization begins.

2. AI cost visibility
Next, you need a reliable view of how much AI you're consuming and what's driving the cost. A monthly spend figure won't get you very far on its own.

Useful visibility breaks consumption down by model, input and output tokens, cache reads and writes, inference type, application, IAM principal, and region. It also tracks cost per million tokens and how usage trends over time.

AWS's CUDOS dashboard already includes AI and machine learning cost analysis, with Amazon Bedrock views showing cost per million tokens by model and by IAM Principal Tag. In August 2026, AWS added Bedrock token usage and prompt caching efficiency visuals as well.

More dashboards aren't the point, though. What matters is making shifts in AI consumption visible enough that FinOps and engineering teams can actually act on them.

3. AI cost governance
Finally, you need controls around that consumption. That means budgets, cost anomaly detection, usage alerts, architectural guardrails, model access policies, clear workload ownership, and spending thresholds.

AI workloads can scale remarkably quickly. Good governance means unexpected usage gets caught early, before it turns up as an unpleasant surprise on the AWS bill.

With these three foundations in place, you're ready to start measuring AI efficiency properly.

The three-layer framework for AI cost optimization
Yuriy's framework splits AI efficiency into three layers:

Technical cost efficiency
Infrastructure unit cost
Business unit cost and value
Each layer answers a different question. Taken together, they give you a far more useful picture than total AI spend ever could.

Layer 1: Technical AI cost efficiency
The first layer asks whether you're making the most of the cost optimization mechanisms already available to you.

These are mostly engineering metrics. The key shift is to stop treating individual optimization techniques as one-off projects and start tracking them as KPIs.

Prompt caching ratio
Prompt caching is a good place to start. Many generative AI applications send the same large chunk of context to a model over and over again. That might be a system prompt, product documentation, a codebase, policy documents, long reference material, or agent instructions. Reprocessing all of it on every request is wasteful.

Amazon Bedrock prompt caching lets supported models reuse prompt content they've already processed, cutting both latency and input token costs.

Pricing varies by model. AWS documentation says that for some supported models, tokens read from cache can be billed at a 90% discount compared with uncached input tokens. Cache writes may be priced separately, though, so it's worth working out the economics workload by workload rather than assuming every cache operation saves the same amount.

Once you think of caching as something to measure rather than just a feature to switch on, you can track it directly:

Prompt caching ratio = cached input tokens ÷ total eligible input tokens

Now the optimization is observable. If one application has a high cache-read ratio while a similar one barely reuses its cache at all, your engineers have a clear lead to follow.

Batch inference utilization
The same thinking applies to batch inference. Not every AI request needs an instant answer. Document classification, large-scale summarization, offline evaluations, content generation, and dataset processing can often run asynchronously instead.

For supported Amazon Bedrock models, AWS prices batch inference at 50% of on-demand inference rates. A useful KPI here might be:

Batch utilization = eligible non-real-time inference processed through batch ÷ total eligible inference

The wider lesson goes beyond these two examples. For every meaningful optimization opportunity, ask whether you can turn it into an efficiency KPI. Once you can, good practice in one team can be measured, repeated, and rolled out across the organization.

Layer 2: Cost per million tokens
Technical KPIs help explain why something is efficient, but they don't always show the overall effect. That's the job of the second layer.

For token-based workloads, cost per million tokens is a particularly handy measure of infrastructure unit cost. It shouldn't be read as an absolute measure of AI value, though. Models differ in how they tokenize, what they can do, how they're priced, and how they perform, and a more expensive model can easily deliver a better commercial result than a cheaper one.

Where cost per million tokens really earns its keep is as a directional signal. Say an application was running at $8 per million tokens in January and $11 by June. That doesn't automatically mean something has gone wrong, but it does tell you something has changed.

Perhaps prompt caching rates have dropped, or outputs are running longer. The application might have switched models, more requests might be using advanced reasoning, or batch workloads might have drifted back to on-demand inference. Or the application could simply be behaving differently than it used to.

Think of it as taking the temperature of an AI workload. A high reading doesn't give you the diagnosis, but it does tell you where to look. From there, you can go back to Layer 1 and work out what's driving the change.

AWS has built this kind of analysis directly into CUDOS, including Amazon Bedrock cost-per-million-token views by model.

Layer 3: Connect AI costs to business value
This is where AI FinOps gets genuinely interesting. The first two layers are mostly about how efficiently you're consuming infrastructure. The third asks the question leadership actually cares about: what are we getting in return?

Answering it means tying AI spend to business outcomes, and the right outcome depends on what the application is for. It might be cost per customer interaction, per support ticket resolved, per document processed, or per software feature delivered. It might be employee hours saved, support tickets avoided, revenue generated, customer conversion, operational tasks automated, or faster time to resolution.

These metrics will look different from one organization, and one workload, to the next. That's exactly as it should be, because there's no universal "AI ROI" metric. AI starts to make economic sense when you can link its infrastructure costs to the specific outcome the application exists to deliver.

The "one more dollar" test
One of the simplest ideas in the framework is what Yuriy calls the one more dollar test. You ask:

Does one more dollar of AI spend generate more than one dollar of outcomes?
It isn't meant to be a precise accounting formula for every use case. Its value is that it forces teams to think about marginal economics.

If extra AI consumption produces disproportionately more value, scaling the workload may well make sense. If costs are climbing faster than the value being created, it's time to dig into the lower layers. The application might need better prompt caching, a different model, or shorter and better-managed context. It might benefit from more batch processing, routing between cheaper and more capable models, architectural changes, or tighter usage controls.

Sometimes the honest answer is that the workload simply doesn't justify its cost. That's a useful conclusion too. AI cost optimization isn't about making every AI workload cheaper. It's about making sure the money you spend produces something worthwhile.

Avoid AI vanity metrics
Generative AI throws off an enormous amount of measurable data. That doesn't mean you should measure all of it.

A good rule of thumb from Yuriy's framework: if a metric can't drive an action or a decision, ask yourself why you're tracking it.

Number of prompts, total tokens, total requests, and number of users can all be useful in context. On their own, though, they slide easily into vanity metrics.

The KPIs worth your attention are the ones that tell you whether your AI architecture is getting more efficient, why your unit cost went up, and which application caused it. They show whether teams are using the optimization mechanisms available to them, and whether extra AI spend is producing extra business value. Those are the numbers that change behavior.

Four steps to improve AI cost efficiency on AWS
If you're trying to bring stronger generative AI FinOps into your organization, the framework boils down to four practical steps.

1. Build the data foundations
Get allocation, visibility, and governance in place, so that AI spend can be traced back to the applications, teams, users, and workloads responsible for it.

2. Score workloads across all three layers
Look at technical efficiency, then infrastructure unit cost, then business value. Don't stop at total spend.

3. Test the economics before you scale
Before significantly expanding an AI workload, check that more spend will actually produce more value. The one more dollar test is a simple way in.

4. Cut the vanity metrics
Prioritize KPIs that lead to action. If nobody would make a different decision when a metric moves, it probably doesn't belong in your FinOps framework.

Where CUDOS fits into AI FinOps
You don't have to build all of this visibility from scratch. AWS's open-source Cloud Intelligence Dashboards, CUDOS among them, make a good starting point.

CUDOS draws on AWS Cost and Usage Report data to give you granular cost and usage analysis. It works with both CUR 2.0 and legacy CUR data, although some of the newer Bedrock attribution features specifically need CUR 2.0.

Its AI and machine learning views currently cover Amazon Bedrock spend, cost per million tokens, IAM Principal Tag analysis, token consumption, and prompt caching insights. That gives you much of what you need for Layers 1 and 2.

Layer 3 usually calls for something more specific to your organization, because that's where AWS billing and usage data has to meet your own application and business data.

AI cost optimization is becoming a FinOps discipline of its own
None of this replaces traditional AWS cost optimization. You still need to manage compute, storage, databases, Savings Plans, Reserved Instances, and the rest of your cloud estate.

But generative AI brings a new set of economics. Consumption can grow incredibly fast. Model choice can swing costs significantly. Agentic architectures multiply requests. And picking the cheapest model won't necessarily give you the most efficient outcome.

That's why the next generation of FinOps needs to connect three things: technical efficiency, unit economics, and business value.

The organizations that get this right won't necessarily be the ones spending the least on AI. They'll be the ones that know exactly where their AI spend is going, why it's changing, and what they're getting back for it.

How Strategic Blue can help
For organizations running significant workloads on AWS, gaining visibility is only part of the challenge.

Strategic Blue helps organizations understand, manage, and optimize their AWS spend, bringing together cloud financial expertise, commercial optimization, and detailed cost visibility.

As Amazon Bedrock and other AI services take up a bigger share of the AWS bill, the same principles matter more and more: understand your consumption, attribute it correctly, find the optimization opportunities, and connect cloud spend to the value it creates.

If your Amazon Bedrock or wider AWS AI costs are starting to scale, talk to Strategic Blue about improving cost visibility and building a more effective approach to AI cost optimization.

Frequently asked questions
How can organizations reduce Amazon Bedrock costs?
There's no single method, because the right approach depends on how the workload behaves. Common places to look include prompt caching, batch inference for eligible asynchronous workloads, model selection, input and output token consumption, application architecture, and cost allocation. What matters most is checking whether each change improves overall unit economics, rather than fixating on token price alone.

Why should FinOps teams track cost per million tokens?
It helps you spot changes in the underlying economics of an AI workload. It's most valuable as a trend rather than a standalone benchmark. A sudden shift is a cue to look at model choice, caching, output behavior, inference type, or another technical driver.

What is different about FinOps for generative AI?
Most traditional FinOps principles still apply, including allocation, visibility, governance, and unit economics. Generative AI adds new variables, such as token consumption, model selection, caching, reasoning behavior, and agentic workflows. It also makes it more important to connect infrastructure consumption directly to business outcomes.

Can CUDOS track Amazon Bedrock costs?
Yes. The current CUDOS dashboard includes Amazon Bedrock cost and usage analysis, with cost-per-million-token views, IAM Principal Tag analysis, token consumption, and prompt caching insights. Some attribution features require CUR 2.0.

[gtm]