Your dashboard says 200ms. Everyone's happy. Support is not.
Averages hide the people you're losing. One request at 30 seconds, ninety-nine at 100ms — the average barely moves. The one person who waited half a minute never comes back.
So measure the ranks instead:
| Percentile | What it describes |
|---|---|
| p50 | The typical experience |
| p95 | A bad day |
| p99 | Your angriest customer |
Now the trap. A page that makes 20 backend calls doesn't get p99 once — it gets 20 chances to hit it. Roughly 1 in 5 page loads contains a p99 call.
Your slowest dependency becomes your typical experience.
This is tail latency amplification, and it's why fan-out architectures feel slower than their numbers suggest.
