The Cloud Outage Your Force-Majeure Clause Already Excludes
The multi-region failure most capable of taking your business dark is deliberate hostile action. It is also the exact scenario your provider's contract, its SLA and your interruption policy are each built to refuse. That gap is uninsured, unpriced, and sitting on your balance sheet.
Open your cloud provider's master agreement and find the force-majeure clause. The AWS Customer Agreement puts events beyond reasonable control outside its obligations, and lists war and terrorism among them. That single sentence quietly decides who pays when everything breaks at once, and the answer is not the provider. The outage most capable of taking out several regions in the same hour, deliberate hostile action, is the exact category the contract carves out. Read your own wording; the convention is near universal. So the failure with the highest correlated impact is the one your provider has already told you it will not fund.
The service-level agreement does not fill that gap, and its own numbers say so. AWS's Compute SLA pays a service credit of 10% once monthly uptime slips below 99.99%, 25% below 99%, and 100% only when it falls under 95%, and every tier is a credit against that service's bill, not cash for lost trade. Do the arithmetic. On a five-figure monthly spend, the maximum payout for a catastrophic multi-day outage is a five-figure credit you can only spend with the provider that just failed you. The SLA was never continuity cover. It is a quality-of-service rebate, and treating it as a financial backstop is the first mispricing.
That leaves insurance, and this is where the gap widens rather than closes. Business-interruption and cyber policies carry war and hostile-action exclusions as standard, and insurers reach for them when the loss is large. NotPetya, the 2017 malware that gutted operations at a pharmaceutical multinational and cost it well into ten figures, became the landmark case: the claim was denied under a hostile-and-warlike-action exclusion, and the fight ran for years. The market's answer since has been to harden, not soften, with Lloyd's moving to require standalone cyber policies to exclude state-backed attacks outright. So the three instruments a board assumes will absorb a correlated outage, the provider contract, the SLA and the insurance policy, are each engineered to step aside from precisely the hostile-action scenario that causes one. I am describing the general market convention here, not adjudicating your specific terms, and you should read your own. But if the correlated outage is the uninsured outage, the exposure does not disappear. It just defaults quietly onto you.
What does an uninsured multi-region outage actually cost?
Take a firm turning over £60m a year. That is roughly £165,000 of revenue riding on a normal trading day. Suppose a correlated event takes its primary provider's European regions down for three days, and suppose, as the wording above allows, both the SLA and the interruption cover treat the cause as excluded. The direct revenue at risk is around half a million pounds, before you count contractual penalties, before goodwill, before the customers who quietly try a competitor while you are dark. These figures are illustrative, put your own in, but the structure holds: the loss is uncapped and it lands entirely on your balance sheet, because the instruments you assumed would absorb it have all stepped aside.
Why does the redundancy you paid for not remove this risk?
Because redundancy is priced for accidents, and a targeted strike is not an accident. Multi-region architecture rests on one assumption: failures are independent. The chance that two regions fall in the same window is near enough the product of two small numbers, which is a much smaller number, so spreading the workload makes a total outage vanishingly unlikely. Against hardware faults, power loss, fumbled config and weather, that maths is sound. An adversary breaks it in a single move, choosing both your Frankfurt and Dublin regions in the same hour because it has read the same architecture diagram you drew. Once failure is chosen rather than drawn, independence is void, and the compound-probability comfort blanket goes with it. The information a customer uses to design failover is broadly the information an attacker uses to defeat it, and providers publish the general shape by design so you can architect against it.
Why is cloud concentration risk a business continuity problem now?
Two shifts moved this from tail risk into the planning horizon. The estate got smaller: Synergy Research Group's public dataset shows the three largest providers account for roughly two-thirds of worldwide cloud infrastructure revenue, so most enterprise capacity now sits inside a handful of provider estates, which is what turns a single hostile act into a correlated, multi-tenant event. The political framing shifted with it: the current US administration casts its AI Action Plan as a strategy to hold American technological dominance, which recategorises commercial data centres as strategic assets in the eyes of anyone opposed to that government. You do not need a named incident to price this; the concentration and the targeting logic are enough on their own. Incidents merely make it concrete, and there have already been claims, to be treated with the scepticism a wartime boast earns, of commercial data infrastructure being named as a target: Iran's Revolutionary Guard, for one, claimed a strike on Amazon data infrastructure in Bahrain, a statement from a party to the conflict with no confirming response reported. Even unproven, it tells you someone now regards a commercial data hub as a legitimate military target. The buildings did not change. Their fencing, their contracts and their insurance did not change. Only the target list did.
Can you plan for the outage you cannot insure?
You cannot buy this risk away inside one estate, and you cannot insure the part that matters, so the mitigation has to be operational. Stop asking whether your provider is resilient. That answer depends on someone else's site security and someone else's intentions. Ask the question you own: what can this business still deliver at 90%, 50% and 10% of capacity, for how long, and at what cost to customers and reputation? Most continuity plans are binary. Up or down. "Fail over to another region" is an assumption wearing a plan's clothes. Graded degradation is cheap to specify, testable, and it survives contact with a correlated failure that no cheque can prevent.
Deciding in advance which services you shed first, which you hold at any cost, and what "running at 10%" concretely means for order intake or safety, is the same discipline as designing systems that stay controllable under stress, and it belongs alongside your technical strategy, not buried in an annexe nobody rehearses. Multi-cloud earns its cost for the handful of must-survive workloads, because it removes the single-estate correlation an attacker exploits, but it buys real complexity too. Treat it as targeted insurance for what has to survive, not a blanket rebuild.
The cloud is not the problem, and the providers built genuinely good machines for the failure modes they were asked to solve. The gap sits in your own plan: the clause you never read to the end, the SLA you mistook for cover, and the exposure you never priced. Find out what your force-majeure wording and your interruption policy actually exclude, then decide what you can still run when the excluded thing happens. That is work to do before the next build, not after the next headline.
Questions people ask
Will my SLA or business-interruption insurance pay out after a deliberate attack on a data centre?
Often not. Cloud force-majeure clauses list war and terrorism as events beyond the provider's obligations, and cyber and business-interruption policies carry hostile-action exclusions that insurers invoke on large losses, as the NotPetya litigation showed. SLA service credits only refund a capped percentage of your bill for the affected service, never lost trade. Read your specific provider agreement and policy wording rather than assuming the scenario that hurts most is covered.
How do I estimate my uninsured cloud concentration exposure?
Start with daily revenue dependent on the affected systems, multiply by a realistic outage duration for a correlated event, then add contractual penalties and a reasoned estimate of customer churn and reputational cost. Subtract only what your policy and SLA will actually pay under an excluded cause, which is frequently close to nothing. The remainder is the figure sitting unpriced on your balance sheet.
How do I build a graded degradation plan for cloud outages?
Define what the business delivers at 90%, 50% and 10% of capacity, decide which services you shed and which you hold at any cost, and cost the customer and reputational consequences of each tier. Then rehearse it. The point is to convert a binary up-or-down assumption into tested decisions you can execute under pressure, since this is the mitigation you can control when insurance and SLAs step aside.
Related
- The Sovereignty Premium: Why Sovereign AI Solutions for Enterprise Are Winning on Access, Not Speed
- You Can't See the Camera Any More: Rewriting Smart Glasses Policy for the Workplace
- AI Safety Policy Is a Moat, Not a Brake
- Security & Trust
Written by an AI editorial persona of Abyshire's proprietary editorial system and reviewed by our team.