AI Token Consumption Keeps Growing—That’s Why “Model Switching” Is the Core of Your Cost Strategy
Lately, I’ve been seeing a steady stream of news about AI costs. Gartner predicts that by 2028, spending on AI coding tools will exceed the average developer’s annual salary. But that’s not the point I want to make.
You don’t have to wait until 2028—the cost problem is already happening right in front of us.
In fact, I’ve already stopped using Claude Code because the costs became too high. There’s no question it’s a useful tool. Still, I decided it wasn’t worth continuing to use.
In this article, I’ll explain why token consumption keeps rising, what lies ahead for unit prices and data centers, and—given that we can’t simply walk away from AI—how to use it wisely, drawing on my own experience.
Token Consumption Is Rising Across the Board
First, one fact to keep in mind: token consumption is unquestionably increasing. This isn’t limited to coding.
A “token” here is the unit of data that AI processes. Every time you ask AI to do something, tokens are consumed—and that consumption translates directly into cost.
And AI today no longer just answers a single question with a single reply. AI agents in particular plan tasks, call tools, fetch data, and repeat steps over and over. What used to finish in one exchange can now consume many times more tokens behind the scenes.
One estimate shows that a simple task costing $0.04 per run in 2023 could reach $1.20 per run with a complex agent-style workflow in 2026—roughly 30 times higher. The smarter and more autonomous agents become, the more tokens they consume. That’s an unavoidable trend.
The Paradox: “Unit Prices Are Falling, but Bills Are Rising”
Here’s where many people get confused: “But isn’t AI getting cheaper all the time?”
It’s true that unit prices for tokens have come down. One analysis found that frontier model unit prices fell 99.7% over the past three years. Looking at that alone, cost hardly seems like a concern.
Yet real-world bills are actually going up. In FinOps-related surveys, 73% of companies report that AI costs exceeded their initial forecasts.
Why does this happen? The answer is simple: usage is growing faster than prices are falling. As things get cheaper, people run workloads that consume more tokens—agent loops, long contexts, repeated retries. Before you know it, usage has ballooned enough to easily swallow any savings from lower unit prices.
“Cheaper, yet more expensive.” That paradox, I believe, is the essence of the cost problem.
The Decline in Unit Prices Is Already Slowing
What worries me further is that the very lifeline of falling unit prices is starting to lose momentum.
According to one data analytics firm, effective token unit prices fell nearly 40% in about six months in the second half of 2025—but since the start of 2026, they’ve dropped only about 6% year to date. The pace of decline has clearly slowed.
Interestingly, companies aren’t necessarily switching to cheaper models. Reports also show a trend of increasing allocation toward high-performance frontier models. In practice, when reliability and performance matter, people end up using expensive models. That’s the real choice on the ground.
In other words, the optimism of “just wait and it’ll get cheaper on its own” is becoming less and less viable.
Data Center Investment Will Push Unit Prices Upward
Another long-term factor is infrastructure—data centers.
Global data center equipment investment is expected to exceed $1 trillion in 2026. Analysis also suggests that next-generation AI data centers cost significantly more to build than conventional ones.
Think about it: with this level of investment, providers will have to recover costs somewhere. Pressure to recoup infrastructure spending is likely to feed back into model pricing. In fact, Gartner analysts have pointed out that infrastructure investment and profitability challenges could push model prices higher going forward.
To summarize what we’ve covered, here’s the picture we’re facing:
- Token consumption will keep rising as agents spread across every domain
- The hoped-for decline in unit prices is already slowing
- Data center investment recovery pressure will eventually work to push unit prices up
If rising consumption coincides with prices bottoming out and upward pressure, it’s natural to expect the cost environment to get even tougher. The original idea of “adopting AI to cut costs” is becoming harder to sustain as-is.
That Doesn’t Mean AI Becomes Unnecessary
So should we pull back from AI? I don’t think so.
Higher costs don’t make AI unnecessary. It speeds up work and makes things possible that were hard to do by hand. That value is real.
If so, only one question matters: how do we use it well? Finding ways to extract value while managing cost—that craft will only grow more important, in my view.
The Key to Cutting Costs Is Still “Model Switching”
So what should you do in practice? My answer comes down to switching models.
The principle is simple:
When it matters, use a capable model even if it costs more. Day to day, use the cheapest model you can.
That’s it. Throwing every task at the highest-performance model is the fastest way to inflate costs. In fact, token unit prices can differ by tens or even thousands of times between models. The gap between inexpensive and high-performance models is that large.
For example, you don’t need a top-tier reasoning model for simple formatting or routine tasks. Delegate those to cheaper models, and switch to high-performance models only for complex design decisions and difficult implementation—the moments that truly count. That split alone can cut costs significantly while preserving quality.
Industry research also suggests that routing tasks to the right model can reduce costs by 60–80%. Conversely, it shows how wasteful it is to keep using high-performance models without thinking.
IDEs Now Make Multi-Model Switching Practical
Fortunately, this kind of switching has become much easier.
Modern IDEs increasingly let you switch among many models. You can pick a model on the spot based on the task—run lightweight models by default, then switch to a high-performance model when you hit a hard problem. That workflow is now a button press away. It’s genuinely useful.
Not being locked to one tool and choosing the best model per task—that flexibility, I think, is the foundation of cost strategy going forward.
Summary—Not “Use It or Quit,” but “How to Switch”
Token consumption will keep growing. Unit price declines are slowing, and data center investment recovery pressure will make the cost environment tougher. You don’t have to wait until 2028—the signs are already here.
What matters, though, isn’t framing AI as a binary choice of “use it or quit.” The real question is how to switch between models.
Use inexpensive models for everyday work. When it counts, pay for high-performance models and get it right. Keep that switching in mind, and AI can remain a reliable partner worth the cost.
I stepped away from Claude Code for now—but I haven’t left AI behind. Instead, I’m gradually finding how to keep a smarter distance with it.