Sold Out: AI Capacity Planning for Enterprise Buyers Starts With Being Told No
Capacity is now a priced product attribute, sold separately from tokens and benchmark rank. The buying decision is arithmetic: what the reserved tier costs, set against what a refusal costs, workload by workload.
On 20 July 2026, a vendor holding one of the most capable models on the market stopped selling it. New subscriptions were paused days after a flagship release, the company saying demand across the previous 48 hours had pushed it close to its capacity limits, and existing subscribers were told they would be prioritised while new places reopened in batches. That is AI capacity planning for enterprise buyers compressed into one announcement: somebody else decides whether you are served, and incumbents go first.
The announcement itself tells you almost nothing about the supplier. Runaway demand produces a pause. So does a compute base thin enough that a routine launch spike swamps it. The observable event is identical, the two causes point in opposite directions, and nobody outside can separate them, because separating them needs utilisation and headroom figures no vendor publishes. As evidence about the supplier, "sold out" and "adding capacity as fast as we can" are marketing artefacts. As disclosures about your position in its queue, they are precise.
The mechanism is dull, which is why it will keep happening. Inference capacity is a physical stock: chips, power, cooling, all with procurement lead times measured in quarters. Attention is a flow that moves in hours. The model in question topped the front-end coding chart on a public evaluation arena after release, exactly the kind of event that turns curiosity into sustained load overnight, and Omdia chief analyst Lian Jye Su put the shortfall down to a chip base that had not anticipated the model's popularity. A stock cannot chase a flow. Any vendor whose product can go viral will eventually ration something.
What does reserved AI capacity actually cost you?
Treat the premium as a number rather than a posture. Take one workload class: a synchronous customer-facing assistant handling two million requests a month, roughly 1,500 input and 400 output tokens each, so about 3.8 billion tokens. Substitute your own rate card, because the arithmetic is the point, not the inputs. At a nominal £2 per million tokens on the rationed on-demand tier, that workload runs about £7,600 a month. Reserved capacity at twice on-demand runs about £15,200. The premium is £7,600 a month, or £91,000 a year.
Now price what the premium buys. Two million requests a month is roughly 2,700 an hour. A four-hour refusal window during a peak therefore pushes about 11,000 conversations onto whatever sits behind the model: human agents, a queue, or an apology. At £4 of incremental handling per conversation, that single window costs £44,000. The annual premium buys out about two such windows. If you expect more than two, reserved capacity is cheap at the price. If you expect fewer, you are insuring against an event that costs less than the insurance, and the honest move is to take the queue.
Run the identical model over an overnight enrichment job with the same token volume and the answer inverts. A four-hour refusal costs a late dashboard. The premium is £91,000 a year for nothing. Same tokens, same vendor, same benchmark score, opposite procurement decision, and the variable that flipped it has no relationship to model quality.
Two forces move that calculation, and neither is the leaderboard. Falling token prices shrink the premium while the cost of a refusal stays denominated in human handling and lost conversions, so cheaper inference quietly widens the set of workloads worth reserving. And the premium exists for a reason: BloombergNEF has data-centre company capex approaching $750bn in 2026 with more than 23GW under construction, against offtake contracts short relative to asset lives. Someone has to service that capital. Guaranteed availability is where it reaches your invoice.
What does AI capacity planning for enterprise actually involve?
The market already sells the thing most buyers assume they have. OpenAI's Scale Tier sells reserved capacity and specified token throughput, and requests beyond your purchased and standard limits can be refused outright with a 429. Assured throughput is a product with a price list. If you have not bought it, you are on the queue, and the queue has a policy you did not negotiate.
The build from there is unglamorous and familiar from every other commodity input. Classify workloads by whether they may wait. Overnight enrichment, backfills, document processing and evaluation runs can queue for hours without anyone noticing. A synchronous customer-facing path cannot queue for ninety seconds. Route the queueable volume to the cheap rationed tier, keep a contracted path on the expensive one, and make failover a designed behaviour rather than an incident. Microsoft's Well-Architected guidance on planned failover for capacity constraints has said the structural version of this for years; the new part is that the constrained resource is a model endpoint. If your architecture cannot express "degrade this workload", that is the gap to close before the next procurement cycle, and it is the same gap our work on AI readiness before you build keeps finding first.
Designed degradation means deciding in advance what the smaller model, the cached answer or the human queue does when the primary is unavailable. Those fallbacks need their own evaluation thresholds, because a silent quality drop during a supply squeeze is worse than a visible delay. Systems built with explicit human control points already have somewhere for the load to go.
Why benchmark rank is the wrong thing to hard-wire
The vendor that could not sell access in July was, that same week, sitting at the top of a public coding chart. Rank and availability are separate variables, and buyers keep purchasing the first while assuming the second comes attached. Leadership positions in this market have changed hands on a cycle measured in months, while platform migrations run in years. Pinning architecture to a leaderboard position is a bet on the fastest-moving variable in the stack.
Dual sourcing is not free, and the case for it is narrower than it sounds. Two suppliers means two evaluation harnesses, two prompt and tool-schema dialects that drift apart, two sets of data-handling review, and a second integration that will rot quietly if nothing routes to it. For plenty of internal workloads the honest answer is to accept the queue and skip the second vendor. The rule that survives scrutiny is the one the cost model produces: dual source the paths where being refused costs more than the redundancy, and contract for priority everywhere else. Working out which paths those are is a technical strategy question with a number attached, not a taste question.
Availability is a feature. You are either paying for it or queueing for it. Price both, per workload, before the next launch surge prices them for you.
Questions people ask
How can I tell whether an AI provider will ration my access during a spike?
You cannot tell from public statements, because a capacity pause looks the same whether it comes from extraordinary demand or from thin headroom. What you can do is ask for the numbers that would settle it: guaranteed throughput, the policy that governs prioritisation between existing and new customers, and the notice period before limits change. A supplier that will not commit to those in writing has answered the question.
What should a capacity clause in an AI supply contract cover?
Reserved throughput expressed in tokens or requests per minute rather than a vague uptime percentage, the behaviour when you exceed it (queued, throttled or refused), your priority relative to other customer classes during constrained periods, notice requirements for changes to rate limits, and whether reserved capacity is portable across model versions when the vendor deprecates one.
Does self-hosting an open-weight model remove the availability risk?
It moves the risk rather than removing it. You stop queueing behind a vendor's other customers and start competing for accelerators, power and colocation on lead times you now own. Self-hosting is a good answer for steady predictable volume with known peaks, and a poor one for workloads whose load can multiply overnight, which is precisely when hosted reserved capacity earns its premium.
Related
- The Memory Oligopoly Behind AI's Cost Inflation
- Least privilege for AI agents is a 1975 idea your pilot is skipping
- The AI Subsidy Cliff: Your Vendor's Investors Have Been Paying Your Bill
- AI & Automation
Written by an AI editorial persona of Abyshire's proprietary editorial system and reviewed by our team.