EN FR ES PT DE AR 中文

AI Factory Cooling Standards Are the Next Vendor Lock-In Trap

The rebrand from data centre to AI factory is a specification, not a slogan. Buildings are being plumbed to one vendor's thermal and power spec before the standards have settled, and the switching costs have moved from the rack to the concrete.

Listen10 min

The most expensive sentence in infrastructure this decade is a rebrand: 'data centre' is out, 'AI factory' is in. The phrase comes from the chip business. Jensen Huang has spent two years describing buildings like these as 'AI factories' that produce intelligence as a commodity, and his company now publishes an 800-volt DC power architecture under exactly that name. Take the metaphor seriously, because it is an engineering document. A data centre was general-purpose plant: standard racks, standard power, air cooling, any vendor's hardware racked next to a competitor's. A factory is single-purpose plant, tooled for one product line. Once AI factory cooling standards are set by the company selling the chips, the building's role changes. It becomes a vendor-specific peripheral with a 15-year mortgage.

Follow the mechanism: why a factory is not a data centre

Follow the mechanism. Current NVL72-class racks already draw well past 100kW each, and the vendor's published roadmap talks of racks approaching 1MW from 2027. For scale, Uptime Institute's global operator survey finds most racks in the field today averaging under 10kW. That is an order-of-magnitude jump, and air cannot carry the heat away fast enough. So the industry is moving to direct-to-chip liquid cooling: cold plates on the silicon, coolant distribution units in the row, manifolds and quick disconnects under the floor. On the power side, high-voltage DC distribution replaces the AC plumbing data centres have used for decades.

Each choice is defensible engineering. The problem is what they add up to. A building plumbed around one vendor's manifold geometry, coolant chemistry and power envelope has made an assumption about whose hardware it will host for the rest of its working life. Switching suppliers then becomes a construction project rather than a procurement exercise.

The efficiency story sold alongside this shift deserves the same mechanical scrutiny. Vendor materials promise steep cuts in cooling electricity and, under the right conditions, configurations that use no water at all. Perhaps. Ignore the press release and check the metering: ask for facility-level figures, independently verified, at your climate and your utilisation. A benchmark published by the benchmark's vendor is marketing in a lab coat.

Who sets AI factory cooling standards? The seller

Standards are supposed to be the boring bit that arrives before the capital. Railway gauges were fought over first, then the network got built, and the backers of the losing gauge paid for the lesson. With AI factories the sequence has flipped. Interfaces, voltages and coolant topologies are still moving, yet the concrete is already being poured. Deploying capital before the standard settles means you are short an expensive option you never consciously sold.

Notice who holds the pen. Nvidia and its peers are promoting their own cooling and power standards, and the company with the strongest incentive to keep an interface proprietary is the company that profits from the dependency. Interoperability is a feature for buyers and a bug for sellers, so expecting the seller to standardise unprompted is like expecting a casino to post the odds at the door.

The counterweight exists, and it is more developed than the lock-in rhetoric admits. The Open Compute Project publishes open rack and liquid-cooling specifications, including blind-mate quick disconnects and shared manifolds, and in 2024 Nvidia itself contributed the rack designs for its GB200 NVL72 system to the project. Which proves the point from the other direction: sellers open interfaces when buyers with leverage insist. Voluntary generosity is not a procurement strategy; documented leverage is.

The money has already moved

At the model layer, the cheques are staggering. OpenAI has reportedly agreed to pay Oracle $300bn over five years from 2027 for compute infrastructure, according to Wall Street Journal reporting on the deal, a commitment neither company has publicly confirmed. Whether that exact figure lands or not, the direction is plain: enormous, long-dated physical commitments, signed now.

The energy forecasts have moved with the money. The International Energy Agency expects global data centre electricity demand to more than double, from around 415 TWh in 2024 to roughly 945 TWh by 2030, with AI the largest driver, and its chief Fatih Birol now calls AI 'one of the biggest stories in the energy world today'. Power on that scale is not a rounding error you can assume away in a lease negotiation.

At the application layer, the returns aren't keeping pace. Accenture's own reporting describes AI value realisation falling short of expectations, with enterprise adoption sluggish outside digital-native organisations. Put the two together and the asymmetry is stark. Capital is committed at the physical layer on the assumption that value will arrive at the application layer. If the second lags the first, what remains is the plumbing.

Can you retrofit a liquid-cooled AI data centre?

This is the question every facilities director should ask before signing, not after. Technically, almost anything can be retrofitted. Economically, the answer is usually no. You're reinforcing floors, rerouting coolant loops, adding distribution units, renegotiating warranties and insurance, and doing it all around live kit. And the density gap makes the job bigger than it looks: going from a sub-10kW hall to 100kW-plus rows is structural work, not a pipe-fitting weekend. When the original build assumed one vendor's geometry, the retrofit bill routinely approaches new-build money.

That's the stranded asset: a hall that still runs, yet whose next useful life costs more than it will return.

The same logic applies to power. Committing to an 800VDC plant today is a bet that the industry's voltage and connector choices in 2029 will resemble today's. If they do, you look clever. If they don't, you own a very efficient museum.

Where does the heat go?

Waste heat reuse is real engineering. Liquid cooling captures heat at temperatures useful for district heating or greenhouses, which is why the brochures show server halls warming housing estates.

The catch is geography. Reuse needs a heat consumer within piping distance, and planning rules in most markets push large data centres into business parks and rural sites, far from the homes that could use the warmth. Until zoning treats these facilities as heat utilities rather than warehouses, most reuse schemes will stay in the brochure.

Germany shows what happens when policy stops waiting. Its Energy Efficiency Act caps new data centres at a power usage effectiveness of 1.2 from July 2026 and obliges operators to reuse a rising share of their waste heat, from 10 per cent at launch to 20 per cent by 2028. Where district heating already exists, that is a compliance task. Where it doesn't, the mandate creates the demand, and the business-park siting logic suddenly looks expensive. Expect other governments to copy the mechanism.

The refresh cycle widens the mismatch in a different direction. Accelerators turn over every few years while buildings amortise over fifteen, and for the electronic waste that generates there is still no serious policy framework anywhere. Every hardware refresh is an environmental liability booked to nobody in particular. Absent regulation, the liability compounds quietly, and it won't stay unpriced for ever.

The steelman: density is real

Now the contrarian test, applied to my own argument. Some of this lock-in is simply physics. At the top end of rack density, no vendor plot is required to explain liquid cooling: nothing else moves that heat. Hyperscalers signing decade-long deals can absorb stranded-asset risk that would bankrupt a mid-market firm, so for them proprietary plant is a calculated bet with a balance sheet underneath it. The technology deserves no suspicion on those grounds. The trap is the tier below copying the architecture without the balance sheet, on leases that outlast the silicon they were signed to host.

How to buy optionality

None of this argues against building. It argues for building with exit ramps. Four tests before capital moves. First, demand open, published interfaces for quick disconnects, manifolds and power connectors; the Open Compute Project's specifications are the obvious benchmark. Second, write retrofit and repowering rights into leases and warranties while you still have bargaining power. Third, stage capital against measured returns rather than projections, the discipline behind getting AI-ready before you build. Fourth, separate the building decision from the chip decision; they have different half-lives.

Pricing those options is what a proper technical strategy is for: valuing the right to switch before you sell it. On the returns side, the organisations actually realising value are the ones running practical AI under human control, rather than the ones hoarding capacity for its own sake.

A standard written by the seller is a toll booth with a spec sheet attached. Let the industry call these buildings factories if it likes. Just make sure the thing being manufactured is your margin, not your dependency.

Edited by Jonathan Taylor.

Questions people ask

What is vendor lock-in in AI factory cooling?

Vendor lock-in here is physical rather than contractual. When a facility's coolant loops, manifold geometry and power delivery are specified around one supplier's hardware, switching vendors means rebuilding plant rather than swapping servers. The switching cost migrates from the rack, where it is cheap, to the building shell, where it is not. Buyers can reduce the risk by insisting on open, published interface specifications before construction starts; the Open Compute Project's rack and liquid-cooling specifications are the current benchmark.

How much does it cost to retrofit a data centre for liquid cooling?

There is no reliable industry-wide figure, because the bill depends on floor loading, existing pipework, electrical distribution and how much downtime the site can tolerate. The shape of the cost is consistent, though: structural work and coolant distribution dominate. Operator surveys show most existing halls averaging well under 10kW per rack, so moving to liquid-cooled rows of 100kW or more is an order-of-magnitude jump, and when the original build assumed one vendor's geometry, retrofit budgets frequently approach new-build economics. Price the retrofit before you sign the original build, not after.

Should enterprises build their own AI infrastructure or rent from hyperscalers?

Match the commitment to the certainty of the workload. Renting suits experimental or spiky demand, because the hyperscaler carries the stranded-asset risk. Owning or long-leasing makes sense only for sustained, predictable utilisation with measured returns behind it. The mistake is signing decade-long physical commitments on the strength of projected value that has not yet been observed. Stage the capital and let the application layer prove itself first.

Related

Written by an AI editorial persona of Abyshire's proprietary editorial system and reviewed by our team.