EN FR ES PT DE AR 中文

The AI Efficiency Paradox: Why Cheaper AI Breaks the Business Model

Token prices have fallen roughly 600-fold since 2020, and every efficiency gain hands a little more of the vendor's margin to the customer. The engineering is succeeding; the business model is the casualty.

Listen8 min

Start with behaviour, not rhetoric. SpaceX is reported to be generating somewhere between $2 billion and $2.5 billion in monthly revenue by renting out what one estimate puts at 400,000 to 500,000 GPUs to AI developers, Google among them, according to Clouded Judgement's analysis of the compute market. Meta, meanwhile, is reportedly preparing to sell its spare compute capacity as a cloud service. When the most aggressive buyers of AI hardware start acting as landlords, the signal is hard to miss: capacity is no longer scarce, and everyone with a balance sheet knows it.

That signal collides with the comforting story of 2026: cheaper inference means more usage, more usage means more revenue, so efficiency is good for everyone in the chain. The first clause is true. The arithmetic behind the second is where the paradox bites, and business models built on scarcity pricing are first in line. A vendor's margin is the gap between what intelligence costs to produce and what it sells for, and that gap is being squeezed from both directions at once.

The scale of the squeeze is easy to understate. One scenario analysis of a potential AI bubble, drawing on recent arXiv research, points to a roughly 600-fold decline in token prices between 2020 and 2026, with a structural break around May 2024, the point at which competition, rather than hardware progress, became the main driver of falling prices. The public record tells the same story at a smaller scale: Stanford HAI's 2025 AI Index put the cost of querying a GPT-3.5-class model at roughly $20 per million tokens in November 2022 and about $0.07 by October 2024, a fall of more than 280-fold in under two years. That break matters more than the headline number. When cost curves set prices, engineers can stay ahead of them. When competitors set prices, every efficiency gain is handed to the customer almost as soon as it's discovered.

Why does cheaper AI inference hurt valuations?

Because the valuations attached to frontier labs are, in effect, option contracts on durable scarcity. Strip the story to its bones: intelligence stays expensive, the handful of firms that can produce it charge rents, and those rents justify the capital deployed. Every link in that chain weakens as inference commoditises. One market forecast expects frontier-tier pricing to fall below $0.50 per million tokens within two years, with budget tiers emerging near $0.10. Prices like that don't support scarcity rents; they support utility margins, and utility margins don't justify frontier-lab multiples.

The same forecast expects vendors to retreat upstream, towards fine-tuning, custom training and vertical integration, with margin shifting from selling inference to adapting it. But that retreat changes the basis of competition. The upstream business rewards distribution, existing enterprise relationships and integration depth over raw model quality, and on that terrain the integrated giants hold most of the cards. A hyperscaler can treat its model as a feature that sells cloud capacity, office software or advertising. A standalone lab needs the model itself to carry the entire profit and loss account.

There's a further twist that few pricing models include. My read is that governments now treat compute as strategic capacity, so hardware demand can stay buoyant even while commercial returns on that hardware go unproven. The spending figures already show demand decoupling from demonstrated returns: Gartner's March 2025 forecast put worldwide generative AI spending at $644 billion for the year, up 76 per cent on 2024, while noting that buyer expectations were cooling after high failure rates in early proof-of-concept work. That props up the silicon layer and the GPU landlords while doing nothing to restore the model vendors' pricing power. Industrial policy extends the boom at the ends of the stack and postpones the reckoning in the middle of it.

Who wins when intelligence becomes a commodity?

The most plausible endgame is segmentation, not collapse. The same scenario analysis sketches a three-tier market: a small premium frontier that stays expensive and concentrated, a commodity cloud tier that becomes brutally competitive, and a growing local and on-device tier absorbing everyday enterprise workloads. If that's roughly right, and it's my base case, enterprise adoption splits in two. Institutional knowledge, the data a firm would never email to a stranger, migrates towards controlled and often local stacks, which is why secure agentic system design matters more than benchmark scores. Commodity tasks flow to whichever cloud is cheapest that quarter.

That bifurcation is a strategy question before it's a procurement one. Firms that sort their workloads early, sensitive versus commodity, will negotiate from strength in both directions. Firms that treat AI as a single undifferentiated service will overpay at the premium end and under-protect at the sensitive end. Getting the sorting right is unglamorous technical strategy work, and it pays for itself the first time a vendor's pricing page changes.

The casualty in that structure is specific: not the hyperscalers, who monetise the glut through rental and bundling; not the enterprises, who capture the surplus as buyers. The loser is the standalone model vendor, squeezed from below by falling prices and from above by the migration of margin into integration.

The second-order consequence almost nobody prices: if the premium tier thins, the risk of stranded GPU capacity lands on whoever owns the hardware. Meta and SpaceX can rent their way out of a glut. A heavily indebted pure-play can't.

What would change my mind

A position without exit conditions is theology, so here are mine.

I'd revisit this thesis if premium-tier token prices held flat for twelve months while capacity kept doubling, because that would show real pricing power rather than temporary inertia. I'd revisit it if enterprise contracts started featuring multi-year locks at premium prices, because that would mean buyers perceive switching costs I currently think are overstated. And I'd revisit it if a standalone lab proved that fine-tuned deployments create data gravity strong enough to survive a cheaper rival's pitch.

On today's evidence, I'd put the probability of durable pricing power for a standalone lab at roughly one in four: high enough to watch, low enough to matter.

The asymmetry nobody priced

Every efficiency gain of the past two years has transferred surplus from vendor to customer, and there's no mechanism in sight that transfers it back. Jeff Bezos built Amazon's expansion on the maxim 'your margin is my opportunity'; the model vendors are now on the receiving end of it. Buyers hold the asymmetry.

The practical response isn't to wait for the dust to settle but to behave like a surplus-taker now: treat frontier inference as a commodity input, negotiate hard and often, keep the crown-jewel workloads on infrastructure you control, and do the readiness work before the build work so a vendor's margin problem never becomes your operational one.

The market priced scarcity. It's being delivered abundance. For the next two years, the smart position is to be long the buyer.

Questions people ask

What is the AI efficiency paradox?

The AI efficiency paradox is the observation that making machine intelligence cheaper to produce can wreck the economics of the firms that produce it. When competition pushes prices down faster than costs fall, every engineering breakthrough is handed to customers as a price cut, so technical success compresses the very vendor margins that high valuations depend on.

Will falling inference costs cause an AI bubble collapse?

Falling inference costs are a sign of healthy competition, but they stress any valuation built on scarcity pricing. The more likely outcome is segmentation rather than total collapse: a concentrated premium tier, a brutally competitive commodity cloud tier and growing local deployment. The most exposed businesses are standalone model vendors without distribution or integration revenue.

Should enterprises run AI locally or buy it from the cloud?

For most organisations the answer is both. Keep workloads involving institutional knowledge and sensitive data on controlled, often local infrastructure, and send commodity tasks to whichever cloud provider is cheapest that quarter. Treating AI as one undifferentiated service leads to overpaying at the premium end and under-protecting at the sensitive end.

Related

Written by an AI editorial persona of Abyshire's proprietary editorial system and reviewed by our team.