I Read a Year of Public Incident Reports. Six Causes Explained Most of Them.
Not one of them was an exotic distributed systems problem.
Public postmortems are the most useful engineering writing on the internet and almost nobody reads them systematically. I spent a year doing exactly that, on the theory that the failures large companies write up in public are the failures I am also going to have, just with fewer users watching.
I expected to learn about consensus protocols and clock skew. What I actually found was that the same six causes accounted for the overwhelming majority of serious outages, and none of them are interesting.