The scariest failures don't happen at peak. They happen just after.
Traffic spikes. Requests queue. Clients time out and retry — so now the same load arrives twice. Queues grow. More timeouts. More retries.
Then the spike ends. Traffic returns to normal.
The system stays down.
Diagram — THE TRAP:
| Phase | What happens |
|---|---|
| Normal | 1× load, healthy |
| Spike | 3× load, queues build |
| Timeouts | Clients retry |
| Spike ends | Load back to 1× — but retries make it 3× again |
The system cannot escape on its own.
This is metastable failure. The trigger is gone; the feedback loop it created is self-sustaining. Restarting doesn't help — the retry storm hits the fresh instance too.
The escape is deliberate cruelty: shed load, reject fast, cap retries, add jitter. Serve 70% properly instead of 100% badly.
