Traffic triples. You want to serve everyone. So you accept every request.
Queues grow. Latency climbs. Requests that are already doomed keep consuming CPU, memory and database connections — all the way to their timeout, at which point the client has already given up and the work is thrown away.
You served nobody. Slowly.
Load shedding is the opposite instinct: reject early, reject cheaply, and protect the requests you can actually finish.
| Strategy | Outcome |
|---|---|
| Accept everything | 100% served badly, 0% on time |
| Shed 30% | 70% served properly |
The hard part is choosing who gets refused:
→ health checks and payments — never
→ retries before first attempts
→ the analytics endpoint before the checkout
And refuse fast. A 503 in 2ms costs almost nothing. A timeout at 30 seconds costs everything you had.
