Azure's West US Outage Was Caused by the Safety Systems Meant to Prevent It

Microsoft's post-incident review of the 23 July failure is a case study in checks that pass individually and fail together.

3 min read ·

For just under five hours on 23 July, customers could not reliably reach Azure services in the West US region. Microsoft's post-incident review (tracking ID ZJV6-SGG) puts the impact between 14:44 and 19:41 UTC, and describes it precisely: traffic entering or leaving the region failed or slowed, while traffic that stayed entirely inside the region was unaffected. The list of affected services runs from App Service and AKS to Cosmos DB, Azure Database for PostgreSQL, ExpressRoute, VPN Gateway and Microsoft Sentinel, with knock-on effects for Microsoft 365.

Early summaries of Microsoft's preliminary review blamed a bug in maintenance software that removed IP routes from more devices than intended. The full review is more interesting, because almost every component in the failure chain was a reliability mechanism doing something it was not designed to do.

How a single repair took out a region's edge

Responses (2)

Sign in to leave a response.

  • The commercial question nobody asks: how many customers paying for zone redundancy assumed it covered this? The review is clear that it did not.

  • Fakhrul

    The recovery-depends-on-the-network point applies to small shops too. Our VPS rollback script pulls from GitHub, which is no help when the box cannot reach GitHub.

More from Daniel Okoye

Recommended from Horizon