The AI Subsidy Cliff: Your Vendor's Investors Have Been Paying Your Bill
Flat-rate AI was priced to win customers, and somebody else covered the difference. As billing shifts to tokens and renewals bite, enterprise buyers are about to meet the true cost of inference.
Somebody has been paying your AI bill, and it was never your finance team. For three years the industry sold frontier intelligence at prices that had nothing to do with cost, because the person swiping the card wasn't the person using the product. That arrangement has a name in every other market: a subsidy. Subsidies end. The AI subsidy cliff is what enterprise budgets hit when this one does, and the gap between what AI costs and what you've been paying for it is the story of 2026.
The useful question isn't whether vendors reprice, because that's already under way. The useful questions are how far the real numbers sit from the ones on your invoice, and what survives the correction.
How much does AI actually cost to run?
Start with the best measurement going. SemiAnalysis bought every consumer AI tier and ran each to its weekly limit on long-horizon coding and agentic work, and its subscription stress-test produced a number now quoted everywhere: a fully used $200 ChatGPT Pro plan, valued at published API rates, would cost up to $14,000 a month. Seventy-fold between sticker price and list value. Measured against true serving cost, which the same analysis puts near a quarter of list price, a maxed-out user burns roughly $3,500 of compute against $200 of revenue: a 17.5-fold gap on the charitable reading.
There's no scandal in any of that; the design worked as intended. A $200 subscription was a customer-acquisition budget with an API attached, and every heavy user was a small, enthusiastic loss.
Follow the mechanism: who was paying, and why they've stopped
Prices are supposed to carry information about cost. These carried none, because the buyer and the payer were different people. The payer was the capital markets. Sequoia's David Cahn did the arithmetic early. His $200 billion question asked where the revenue to justify AI infrastructure spend would actually come from, and by June 2024 the gap had tripled into his $600 billion question, with capital deployment outrunning every credible revenue path. The scale of that deployment is public: Microsoft alone said it was on track to spend roughly $80 billion on AI datacentres in a single fiscal year.
Free-to-play games run on this economy. A handful of whales fund the lobby, and everyone else plays on their credit cards. Enterprise AI inverted it: the whales were venture funds and big-tech balance sheets, and the whole industry played for free. The difference is that games get to keep their whales. Ours have quarterly reporting, obligations to service and, eventually, public markets to face. Investors will subsidise adoption. They won't subsidise habit.
Seen this way, the chatbot was the packaging and the subsidy was the product. Once you accept that, the year's pricing news stops looking like a series of isolated decisions.
The workload broke the pricing first
Flat-rate pricing survived while the workload was chat: short prompts, short answers, predictable averages. Then the industry shipped agents, and the averages went feral. An agentic coding session plans, reads files, writes patches, runs tests and loops; a single task can burn roughly a thousand times the tokens of a chat answer, and frontier models will spend over 100,000 output tokens before they stop. Vendors understood the maths early, which is why GitHub's own plan pages now price heavy agentic usage in metered 'premium requests', charged per request beyond the monthly allowance. All-you-can-eat pricing is fine when customers order the salad. It breaks when they hire the kitchen.
The first big casualty arrived in April 2026, when Microsoft paused new sign-ups for GitHub Copilot's student and individual tiers after operational costs nearly doubled in a quarter. The public trail runs through GitHub's own Copilot changelog: rate limits tightened, allowances formalised, billing steered onto tokens.
Watch the shape of that response: no press release announcing a rise. When a vendor changes the unit of billing, the unit is the rise. Token billing reprices your heaviest users first and quietly moves cost risk from the vendor's P&L to your budget.
The productivity claims were never metered
The correction says nothing about whether the tools work. It says nobody measured them properly while the compute was free, and the most honest measurement to date is awkward for everyone. When METR ran a randomised trial with experienced open-source developers, AI assistance made them 19% slower while they believed they were 20% faster. One study, one setting, but it remains the only number of its kind, and it shows how little the industry knows about its own product. Vendor benchmarks measure what flatters: pass rates on curated tasks, speed on greenfield code. The number a buyer actually needs, cost per completed task, almost never appears in the sales deck.
A tool that saves a developer an hour while burning £40 of subsidised compute is one purchase. The same tool at true serving cost is a different one, and neither observation tells you the tool is useless. What they tell you is that the business case was written in someone else's money. If a productivity deficit exists, it lives in the accounting before it lives in the models.
Is running your own models cheaper?
The obvious escape hatch is narrower than it looks. Open-weight models cut the per-token bill to zero and add hardware, power and maintenance in its place. Price published on-demand GPU rates against a small team's workload and the API bill usually wins, particularly on the long-context tasks agents need. Self-hosting buys control, data residency and independence from a vendor's pricing mood. It does not buy the subsidy back. If your plan for the cliff is 'we'll just run open weights', price the power bill before you price the freedom.
Where the AI subsidy cliff leaves enterprise budgets
Three consequences follow. None is optional.
First, procurement shifts from fear of missing out to arithmetic. Every renewal in the next eighteen months is a repricing event, so walk in with your own usage data and negotiate on cost per completed task rather than seat price. A benchmark without a cost column is an advert.
Second, efficiency becomes an engineering discipline again. Token budgets, model routing, caching, smaller models for smaller subtasks: the unglamorous crafts the subsidy made optional will now separate teams that meter from teams that guess. If your AI usage has no metering, fixing that is an AI readiness question before it's a tooling question. Agentic systems in particular need governance and observability built in rather than bolted on, which is the entire premise of designing secure agentic systems.
Third, expect consolidation. Vendors whose unit economics close at honest prices will survive, because the demand is real. Vendors whose growth depended on the subsidy get bought, merged or repriced into respectability. No prophecy is required here; the mechanism only points one way.
None of this makes AI a bubble in the usual sense. The models are genuinely useful; the doubt was always attached to the price. For three years the subsidy was the product, the renewal is the invoice, and this is the year they meet. Budget for the model you can afford at cost, not the demo you were shown at a loss. A technical strategy that prices inference honestly will outlast any benchmark report.
Questions people ask
When will AI prices increase for enterprise buyers?
They already are, but through mechanism rather than announcement: flat-rate tiers are being retired, rate limits tightened and billing moved to tokens, which reprices the heaviest users first. Treat every renewal from 2026 onwards as a repricing event and negotiate on cost per completed task, not seat price.
How much does AI actually cost to run at full utilisation?
SemiAnalysis stress-testing found that a fully used $200 ChatGPT Pro plan would cost up to $14,000 a month at published API rates, and roughly $3,500 against estimated true serving cost. The gap between subscription price and compute consumed sits between 17x and 70x depending on how you measure it.
Is self-hosting open-source models cheaper than paying for AI APIs?
Usually not at the frontier. Per-token cost falls to zero, but hardware, electricity and maintenance frequently exceed the API bill, especially for the long-context tasks coding agents need. Self-hosting buys control and data residency; it does not recreate the subsidy.
Related
- The Memory Oligopoly Behind AI's Cost Inflation
- AI Infrastructure's Circular Financing Risk: a Hall of Mirrors With a Mortgage
- AI layoffs 2026: the jobs apocalypse is a sales pitch, and the backlash is the real risk
- AI & Automation
Written by an AI editorial persona of Abyshire's proprietary editorial system and reviewed by our team.