EN FR ES PT DE AR 中文

The AI Capex Bet Is on the Wrong Chips

The infrastructure boom is underwritten by a claim about inference that the economics of inference keep undercutting. Match the workload to the hardware, and a chunk of the capex is insuring against demand its own logic points away from.

Listen7 min

The AI capex bet has a hardware problem that the spending is not pricing in. Hundreds of billions in AI infrastructure spending are being justified by one word, inference, and the assumption underneath it is that serving all this AI to the world will keep demanding the newest, most expensive chips. That assumption is quietly wrong. Training a frontier model needs the latest silicon. Running the models people actually use, most of the time, does not.

Follow the mechanism. There are two AI workloads, and they have almost nothing in common. Training is a one-off capital event, where a small number of labs burn through the fastest accelerators money can buy to produce a model. Inference is the volume business, the billions of small requests that happen once the model exists. Here the economics pull hard in one direction: a small, task-specific model that answers well enough runs on silicon several generations behind the frontier, at a fraction of the cost per token. As frontier model pricing gets cut and inference commodifies, the rational move is to match the model to the task, then serve it on the cheapest chip that clears the bar. No public dataset breaks production inference down by chip generation, so this is an argument from the economics of the workload, and the incentive it describes is not in doubt.

The efficiency numbers push the same way. Stanford HAI's 2025 AI Index documents inference costs falling by orders of magnitude and small open models closing much of the gap on the benchmarks enterprises care about. The work that actually pays the bills, sorting a support queue or summarising a document, rarely needs a trillion-parameter model. It needs a good-enough one that answers in milliseconds and costs almost nothing to serve.

Here is the tell for anyone reading the boom. If usage scales on older, cheaper silicon, then growth in AI usage and growth in frontier-chip demand should come apart, and capex justified by 'inference demand' is buying insurance against a future the economics of the workload keep pointing toward. What would confirm it: disclosures showing inference volumes climbing while frontier-accelerator utilisation flattens. What would refute it: frontier fleets staying saturated as usage rises. Either way, 'usage is up' is not the same claim as 'the frontier tier is used more', and the capex thesis quietly treats the two as one.

Is the AI capex boom building the wrong hardware?

Start with the size of the cheque. Gartner forecast worldwide generative-AI spending of around 644 billion dollars for 2025, and hyperscaler capital budgets have kept climbing since. Against that, Sequoia's 600 billion dollar question asked the awkward part out loud: where is the revenue all this hardware is supposed to earn back? The gap is being papered over by circular financing, where a supplier funds the customer that then buys its product, from chipmakers backing the firms that buy their chips to content owners investing in the model makers that will pay them for licences. Each of those deals books as demand. Not all of it is organic.

None of this means a crash. AI is more likely to burn than crash, a correction that washes out the over-leveraged while leaving the physical plant behind. That is the dark-fibre story, and it is partly right. Analysts already describe operators preparing to resell spare compute as a cloud service, which is what a glut looks like before it clears. But the analogy has a limit the optimists skip. Dark fibre kept its value because a strand of glass from 2001 still moves photons today. A frontier GPU cluster superseded by the next generation is not fibre. It is the part of the buildout with no second life, and it is precisely the tier the capex thesis is most exposed to. Any honest unwind scenario lands the damage unevenly: the reusable plant finds a floor, the superseded frontier clusters do not.

There is a wrinkle that complicates even the cheerful reading, and it is happening now. The plan to run inference on cheap, older hardware assumes cheap hardware stays cheap. Memory says otherwise. TrendForce's March 2026 survey forecast DRAM contract prices rising 58 to 63 per cent quarter on quarter, with NAND up 70 to 75 per cent, and Micron announced it would exit its Crucial consumer memory business to point capacity at data-centre buyers. When the buildout eats the memory supply, the second-hand inference tier gets more expensive too. The frontier's appetite raises the floor under everyone.

What should a business sizing its own AI spend do?

For most companies the answer is not to build to the frontier at all. If your workload is inference, the honest question during any AI-readiness assessment is which of your tasks genuinely needs a frontier model and which are being over-served by one out of habit. Sizing infrastructure to the biggest model you might ever call is how you end up owning depreciating assets that a smaller model on a cheaper chip would have handled. This is a capital-allocation decision before it is a technical one, and it belongs in your technical strategy, not in a procurement reflex to buy the newest thing.

The policy backdrop makes planning harder, not easier. Federal AI rules are unsettled after tech-industry lobbying pushed the administration to cancel a planned AI executive order, while states legislate on their own, from California's SB 53 onward. A company trying to set a stable compliance baseline is aiming at a target several parties are actively moving. The same pressure to monetise is starting to seat advertisers inside decision loops that used to be neutral, which is one more reason to keep a human in control of any AI-assisted call as the money gets involved.

The buildout will leave something behind. Just make sure it is not your budget, tied to the one tier of hardware this logic tells you to buy last.

Questions people ask

Do you need the latest GPUs to run AI inference?

For most production inference, no. Frontier accelerators earn their price during training. Serving small or distilled models for narrow tasks runs acceptably on hardware several generations old, and the falling cost per token rewards matching the model to the task rather than defaulting to the biggest one you can call.

Does cheaper inference mean AI demand is falling?

No, and conflating the two is the mistake. Usage can keep rising while frontier-chip demand flattens, because the extra usage can land on cheaper, older silicon. Falling inference prices are a sign of commodification, not shrinking demand.

Will AI data centres become stranded assets?

Some of the physical plant, the power and cooling a future tenant can reuse, along with inference-grade hardware, is likely to find a second life if operators over-build, much as dark fibre did. The exception is top-end training hardware, which loses value fast when the next chip generation arrives.

Related

Written by an AI editorial persona of Abyshire's proprietary editorial system and reviewed by our team.