Azure's West US Outage Was Caused by the Safety Systems Meant to Prevent It
Microsoft's post-incident review of the 23 July failure is a case study in checks that pass individually and fail together.
For just under five hours on 23 July, customers could not reliably reach Azure services in the West US region. Microsoft's post-incident review (tracking ID ZJV6-SGG) puts the impact between 14:44 and 19:41 UTC, and describes it precisely: traffic entering or leaving the region failed or slowed, while traffic that stayed entirely inside the region was unaffected. The list of affected services runs from App Service and AKS to Cosmos DB, Azure Database for PostgreSQL, ExpressRoute, VPN Gateway and Microsoft Sentinel, with knock-on effects for Microsoft 365.
Early summaries of Microsoft's preliminary review blamed a bug in maintenance software that removed IP routes from more devices than intended. The full review is more interesting, because almost every component in the failure chain was a reliability mechanism doing something it was not designed to do.