Mark Russinovich and the Azure reliability team published a post last month that makes an uncomfortable argument: Microsoft Azure engineers have concluded that your architecture diagram proves nothing about whether your workload can survive a bad Tuesday. A diagram is a statement of intent. Only a test tells you what you actually have.
The post, co-written by Russinovich, Adam Bogobowicz, and Molina Sharma, frames resilience as something most teams set up once and then quietly stop thinking about. Configure disaster recovery, write a runbook, run a failover drill during onboarding, done. That approach kept the lights on in a simpler era, but it treats resilience as a project with an end date instead of a property you maintain for the life of the system.
What a diagram cannot tell you
The part that lands hardest is a short list of questions a diagram structurally cannot answer. It cannot tell you whether your resiliency goal is being met right now. It cannot define what “resilient” means for that particular application. It cannot confirm the failover path actually works today, as opposed to the day someone drew it. And it cannot show you the resources nobody bothered to draw.
That last one is the sneaky failure mode. The health probe pinned to a single subnet, the connection string pointing at the primary region, the encryption keys that live only in the primary key vault. None of that appears on the picture, and all of it decides whether recovery works. The post’s example is a good one: keys that exist only in the region that just went down, discovered missing at the exact moment you need them most.
Change is the real killer
Azure’s own numbers give the reason diagrams drift. Roughly 70 percent of cloud outages relate to change in some way, and usually it is an ordinary modification whose blast radius nobody re-evaluated. Someone rotated a credential. Someone resized a pool. Someone added a dependency. The architecture did not change on paper, so the picture stayed accurate and the system stopped being resilient.
Microsoft’s own response is deployment discipline. Changes roll to a canary region first, then a pilot region with deliberate bake times, before anything touches the broader fleet. The Well-Architected Framework describes this as safe deployment practice, and it exists precisely because resilience is a threshold you can silently fall below.
AI made this worse, not better
The section I did not expect in a resilience post is the one about AI dependencies. The argument: the dependency that breaks a modern workload is increasingly probabilistic. An inference endpoint, a retrieval pipeline, a model that answers differently each call. None of these appear on an architecture diagram, they can throttle or fail without the infrastructure around them looking unhealthy, and they change every time you adjust a prompt or swap a model.
The team’s advice is to treat model changes like any other software release, because that is what they are. Change the prompt, the model, or the harness around it and you have changed the software. Their internal practice is to stay deterministic wherever possible, wrap non-deterministic components in deterministic checks, and use adversarial review agents only where deterministic checks cannot reach. Evaluation against explicit SLOs is the step teams skip, and according to the authors it is the one that matters.
The tooling catching up
Two Azure services back the argument with something runnable. Azure Chaos Studio now ships Workspaces and preconfigured Scenarios, including Compute Zone Down, DNS Outage, and database failover patterns, currently in public preview. The scenario model matters because it lowers the barrier: instead of hand-crafting fault injection experiments, you pick a scenario, scope which resources it can touch, and run it on a cadence.
The other is Azure Infrastructure Resiliency Manager, also in public preview. It is a global non-regional service, free during preview, built around helping teams start resilient, get resilient, and stay resilient. Internally, Azure pairs this with machine learning applied to observed service behavior, so “healthy” is defined by what a service actually does rather than by what the design assumed.
What to actually do this quarter
Strip out the philosophy and the post reduces to a practical checklist. Define explicit resiliency goals and SLOs per application, then set RTO and RPO numbers you have verified rather than numbers you hope are true. Verify the recovery path’s own dependencies, including whether your keys and configuration survive the loss of a region. Run chaos experiments on a schedule: combine load testing with fault injection in the same run, compare degraded-state metrics against a baseline, and keep separate baselines that tolerate expected error spikes. Wire those tests into CI/CD as deployment gates so a resilience regression fails a build instead of causing an incident.
Scope matters when you start injecting faults. The Chaos Studio guidance says to limit the workspace identity so it can only affect what it should, exclude critical resources per scenario, and start in non-production before pointing anything at real users.
One more point from the post worth repeating: resilience is a cost decision as much as a design one. A workload carrying $100M of revenue on a single trading day may justify active-active cross-region topology for that day and active-passive for the rest of the year. Matching the topology to the actual business exposure is the design, not a compromise.
The bottom line from the authors is blunt: a diagram tells you what you intended. Only a test tells you what you have. If your last failover test predates your last three architecture changes, the diagram is the most confident lie in your documentation set.