Cards/Case StudySystem Design · Day 8Aug 8, 2026

The Outage That Starts After You're Saved

The Outage That Starts After You're Saved — system design card, day 8, case study

The scariest failures don't happen at peak. They happen just after.

Traffic spikes. Requests queue. Clients time out and retry — so now the same load arrives twice. Queues grow. More timeouts. More retries.

Then the spike ends. Traffic returns to normal.

The system stays down.

Diagram — THE TRAP:

PhaseWhat happens
Normal1× load, healthy
Spike3× load, queues build
TimeoutsClients retry
Spike endsLoad back to 1× — but retries make it 3× again

The system cannot escape on its own.

This is metastable failure. The trigger is gone; the feedback loop it created is self-sustaining. Restarting doesn't help — the retry storm hits the fresh instance too.

The escape is deliberate cruelty: shed load, reject fast, cap retries, add jitter. Serve 70% properly instead of 100% badly.

Systems don't recover just because the cause went away.