On January 31, 2017, a GitLab engineer ran rm -rf on the PostgreSQL data directory of what he thought was the replica. It was the primary. About 300 GB of GitLab.com's production database was gone within a second or two, and when the team reached for backups, none of the five they had worked. GitLab.com was down for about 18 hours and lost six hours of data for good. It is the most famous database postmortem there is, and nearly every cause in it is still sitting in someone's infrastructure to...