Why did recovery take almost 3 hours in one region when others recovered in 40 minutes?
The us-central1 region hit a 'herd effect': once the kill switch fixed the underlying bug, every crashed Service Control instance tried to restart at the same moment, all hammering the same Spanner database for policy metadata simultaneously. Without randomized exponential backoff on the restarts, the recovery attempt itself became a second, self-inflicted overload.
Answered in
The Google Cloud Outage That Started With One Missing Null CheckOne unprotected code path in Google Cloud's Service Control took down Spotify, Discord, and Cloudflare auth in minutes. Here's the anatomy of the failure.
Read the full analysisOther questions this article answers
More system design questions
- Why does a database need an index at all — why can't it just scan the table?
- When is a hash index better than a B-tree index?
- What is a composite index and why does column order matter?
- What's a covering index and why is it faster than a normal index?
- What's the real cost of adding an index, beyond disk space?
- What does ACID actually stand for, and why do all four properties matter together?
- What's the practical difference between pessimistic and optimistic concurrency control?
- What is a race condition in a database transaction, and how does isolation prevent it?
Every answer on Crashtech is written by the editor of the article it comes from — never auto-summarised. Browse all answers or the System Design beat.