What happens when the cloud blows up?
InfoWorld ·

I’ve covered cloud computing for a long time, and one of the most persistent myths I encounter is that public clouds exist above the laws of physics that govern everything else in technology. The events coming out of the Middle East should put that myth to rest for good. Reuters reported this week that Amazon Web Services says it cannot restore access to its cloud region in Bahrain and to one of its three data-hosting zones in the United Arab Emirates, following damage during the Iran war. The status update is stark reading. AWS admitted that the damage in Bahrain spanned multiple availability zones (AZs) and, in its own words, “exceeded what our regional and multi-AZ services are designed to withstand.” Think about that for a moment. The entire architectural premise of availability zones is that they fail independently, so a multi-AZ deployment survives a localized disaster. The war destroyed that assumption in a single stroke. The disruptions began in March, after US and Israeli attacks on Iran prompted Tehran to retaliate with missile and drone strikes on Israel and on Gulf states hosting US military bases. Two of AWS’s facilities in the UAE were directly struck, and a drone strike close to a Bahrain facility physically damaged infrastructure there. A second Bahrain availability zone went down in April, and the region became effectively unavailable. In the United Arab Emirates, AWS says it cannot restore resources and data hosted exclusively in the zone known as mec1-az2. The company has notified authorities in both countries and is helping customers re-establish operations in other regions. AWS expects to provide a UAE restoration update in the coming months, with a Bahrain update not until early 2027. Read that timeline again: Operations interrupted in March may not be resolved until early next year. That’s a year of impact for data and workloads that weren’t migrated in time. Some of that data, I suspect, will never come back. I bring this up not to pile on AWS. Frankly, no cloud provider could have engineered its way out of a drone strike on the infrastructure surrounding a data center. However, this event reveals something important about how we all think about cloud computing. Public clouds are not flawless Here is the uncomfortable truth too many enterprises still refuse to internalize: Public clouds are not magic. They are data centers just like the data centers we’ve been running for decades. They are filled with servers, storage arrays, switches, fiber, and generators. They sit in buildings, near power grids and undersea cables, in countries subject to changing geopolitics. They are susceptible to everything traditional data centers have always been susceptible to: natural disasters, network outages, power failures, human error, and yes, war. Most people don’t consider this when they architect their systems, and I get why. It feels remote. But here we are, with a hyperscaler publicly confirming that physical damage from military conflict exceeded its own resilience design. If it can happen to AWS in Bahrain, the question isn’t whether it can happen elsewhere. The question is whether your planning accounts for it. Cloud providers, to be fair, have done a pretty good job maintaining uptime over the years. But good is not flawless, and 2025 reminded everyone of that in a very expensive way. Last year, major cloud providers, including AWS and Microsoft, suffered significant outages that cost many companies billions of dollars. Here’s the kicker: Many of the companies absorbing those losses weren’t even users of those clouds. They were customers of a software-as-a-service provider or another infrastructure service provider that happened to use Microsoft or AWS to serve their infrastructure. This idea of a cloud computing supply chain opened a lot of minds last year. Your application might not depend on a hyperscaler directly, but your payroll vendor, your CRM, or your logistics platform might. When the cloud beneath your cloud goes down, you go down with it, and you have no contractual relationship with the thing that actually failed. Many came to realize that the durability of their supply chain of cloud computing was a real issue. Cloud computing is susceptible to outages just like enterprise infrastructure always has been, and pretending otherwise is not a strategy. Build for what you hope won’t happen So, what should enterprises actually do? The answer is not to flee the cloud, and it’s not to accept whatever happens as an act of God. The answer is to consider possible unforeseen disasters and build them explicitly into your business continuity and disaster recovery plans, rather than assuming the provider will handle everything. Here are three things I would urge every enterprise to consider when evaluating cloud resiliency and trying to eliminate as much risk as they reasonably can. First, know your complete dependency chain , including your SaaS providers and infrastructure partners, down to the cloud regions and availability zones they run on. Most enterprises discovered last year that they couldn’t answer the question “Which of my critical services depend on AWS?” without weeks of investigation. If your provider’s data center is destroyed, whether by fire, flood, or a missile, you need to know exactly what’s exposed—and before it happens, not after. Second, design for region loss , not just zone loss. Multi-AZ deployment is table stakes, but Bahrain demonstrates that availability zones are not invulnerable to large-scale physical damage. Genuinely resilient architectures span cloud regions and, in some cases, multiple providers, with data replication and automated failover. Yes, this costs more. Ask yourself what a year of unavailability would cost you, and the math usually changes quickly. Third, test your recovery like you mean it. A plan that has never been exercised is a document, not a capability. Run game days that simulate the loss of an entire region, including the SaaS providers in your supply chain. Verify your backups exist in a location physically and logically independent of the failure domain. Can you actually restore from them within your stated recovery time and recovery point objectives? The enterprises that came through the 2025 outages best were the ones that had rehearsed failure, not the ones with the thickest disaster recovery binders. Until recently, war damage to public cloud infrastructure was something none of us put on our risk registers. Now it has a data point attached, with a restoration timeline measured in many months. The cloud remains the right choice for most workloads, and providers earn their reputation on reliability. But reliability at scale is not the same thing as invulnerability. Clouds are physical, and physical things break. Plan accordingly.
I’ve covered cloud computing for a long time, and one of the most persistent myths I encounter is that public clouds exist above the laws of physics that govern everything else in technology. The events coming out of the Middle East should put that myth to rest for good. Reuters reported this week that Amazon Web Services says it cannot restore access to its cloud region in Bahrain and to one of its three data-hosting zones in the United Arab Emirates, following damage during the Iran war. The status update is stark reading. AWS admitted that the damage in Bahrain spanned multiple availability zones (AZs) and, in its own words, “exceeded what our regional and multi-AZ services are designed to withstand.” Think about that for a moment. The entire architectural premise of availability zones is that they fail independently, so a multi-AZ deployment survives a localized disaster. The war destroyed that assumption in a single stroke. The disruptions began in March, after US and Israeli attacks on Iran prompted Tehran to retaliate with missile and drone strikes on Israel and on Gulf states hosting US military bases. Two of AWS’s facilities in the UAE were directly struck, and a drone strike close to a Bahrain facility physically damaged infrastructure there. A second Bahrain availability zone went down in April, and the region became effectively unavailable. In the United Arab Emirates, AWS says it cannot restore resources and data hosted exclusively in the zone known as mec1-az2. The company has notified authorities in both countries and is helping customers re-establish operations in other regions. AWS expects to provide a UAE restoration update in the coming months, with a Bahrain update not until early 2027. Read that timeline again: Operations interrupted in March may not be resolved until early next year. That’s a year of impact for data and workloads that weren’t migrated in time. Some of that data, I suspect, will never come back. I bring this up not to pile on AWS. Frankly, no cloud provider could have engineered its way out of a drone strike on the infrastructure surrounding a data center. However, this event reveals something important about how we all think about cloud computing. Public clouds are not flawless Here is the uncomfortable truth too many enterprises still refuse to internalize: Public clouds are not magic. They are data centers just like the data centers we’ve been running for decades. They are filled with servers, storage arrays, switches, fiber, and generators. They sit in buildings, near power grids and undersea cables, in countries subject to changing geopolitics. They are susceptible to everything traditional data centers have always been susceptible to: natural disasters, network outages, power failures, human error, and yes, war. Most people don’t consider this when they architect their systems, and I get why. It feels remote. But here we are, with a hyperscaler publicly confirming that physical damage from military conflict exceeded its own resilience design. If it can happen to AWS in Bahrain, the question isn’t whether it can happen elsewhere. The question is whether your planning accounts for it. Cloud providers, to be fair, have done a pretty good job maintaining uptime over the years. But good is not flawless, and 2025 reminded everyone of that in a very expensive way. Last year, major cloud providers, including AWS and Microsoft, suffered significant outages that cost many companies billions of dollars. Here’s the kicker: Many of the companies absorbing those losses weren’t even users of those clouds. They were customers of a software-as-a-service provider or another infrastructure service provider that happened to use Microsoft or AWS to serve their infrastructure. This idea of a cloud computing supply chain opened a lot of minds last year. Your application might not depend on a hyperscaler directly, but your payroll vendor, your CRM, or your logistics platform might. When the cloud beneath your cloud goes down, you go down with it, and you have no contractual relationship with the thing that actually failed. Many came to realize that the durability of their supply chain of cloud computing was a real issue. Cloud computing is susceptible to outages just like enterprise infrastructure always has been, and pretending otherwise is not a strategy. Build for what you hope won’t happen So, what should enterprises actually do? The answer is not to flee the cloud, and it’s not to accept whatever happens as an act of God. The answer is to consider possible unforeseen disasters and build them explicitly into your business continuity and disaster recovery plans, rather than assuming the provider will handle everything. Here are three things I would urge every enterprise to consider when evaluating cloud resiliency and trying to eliminate as much risk as they reasonably can. First, know your complete dependency chain , including your SaaS providers and infrastructure partners, down to the cloud regions and availability zones they run on. Most enterprises discovered last year that they couldn’t answer the question “Which of my critical services depend on AWS?” without weeks of investigation. If your provider’s data center is destroyed, whether by fire, flood, or a missile, you need to know exactly what’s exposed—and before it happens, not after. Second, design for region loss , not just zone loss. Multi-AZ deployment is table stakes, but Bahrain demonstrates that availability zones are not invulnerable to large-scale physical damage. Genuinely resilient architectures span cloud regions and, in some cases, multiple providers, with data replication and automated failover. Yes, this costs more. Ask yourself what a year of unavailability would cost you, and the math usually changes quickly. Third, test your recovery like you mean it. A plan that has never been exercised is a document, not a capability. Run game days that simulate the loss of an entire region, including the SaaS providers in your supply chain. Verify your backups exist in a location physically and logically independent of the failure domain. Can you actually restore from them within your stated recovery time and recovery point objectives? The enterprises that came through the 2025 outages best were the ones that had rehearsed failure, not the ones with the thickest disaster recovery binders. Until recently, war damage to public cloud infrastructure was something none of us put on our risk registers. Now it has a data point attached, with a restoration timeline measured in many months. The cloud remains the right choice for most workloads, and providers earn their reputation on reliability. But reliability at scale is not the same thing as invulnerability. Clouds are physical, and physical things break. Plan accordingly.