These two windows have the same mean and different users.
Imagine each row is one window of 100 requests. In the steady window, every request takes 195 ms. In the second, 95 requests take 100 ms and five take 2,000 ms. Both means are 195 ms: the fast majority offsets the slow requests when all 100 values are added together.
| Window | Mean | Median | p95 | p99 | Variance |
|---|---|---|---|---|---|
| Steady: 100 × 195 ms | 195 ms | 195 ms | 195 ms | 195 ms | 0 ms² |
| Long tail: 95 × 100 ms, 5 × 2,000 ms | 195 ms | 100 ms | 100 ms | 2,000 ms | 171,475 ms² |
The mean is the sum divided by the count, so both windows average to 195 ms. The median is the middle observation after sorting: it shows the typical request shifted from 195 ms to 100 ms. The table also shows why the mean alone cannot tell you whether the slow requests matter to users.
Spread matters when requests do not behave alike.
The range reports only the minimum and maximum. Variance uses every observation: subtract the mean from each value, square each difference, then average those squared differences. Squaring keeps fast and slow deviations from canceling and gives distant values more influence.
For the 100-request long-tail window above, population variance is 171,475 ms². Standard deviation is its square root, about 414 ms,
which brings the spread back to latency units. The steady window has zero variance because every
request takes the same time.
Variance is useful when comparing how dispersed two sets of measurements are, but squared units are hard to explain as a user-facing target. Standard deviation is still sensitive to extreme values and does not say what share of requests crossed a deadline. Use percentiles when the question is about a latency threshold.
p95 is a threshold for most requests, not a report of the worst ones.
For nearest rank, sort n observations, calculate the one-based position ceil(0.95 × n), and select that value. With 100 observations, p95 is item 95.
In the long-tail example, items 1 through 95 are 100 ms, so p95 is 100 ms; p99 reaches the
2,000 ms requests. A p95 tells you where the 95% threshold falls. It does not summarize the
remaining five percent.
Nearest rank is one convention. Some tools interpolate between observations, so compare systems using the same definition. Do not average two reported p95 values to get a combined p95; use the underlying observations or a mergeable distribution summary whose approximation limits are known.
Move the tail and watch each summary respond.
Start with the same-mean example. The two windows both average 195 ms, but their p99 and variance tell a different story. Change one or two current measurements, then compare the median, p95, p99, and maximum. Try the small-sample example too: with only five readings, nearest-rank p95 selects the maximum, so one request controls the result.
Compare two latency windows
Edit either list or paste up to 100 non-negative measurements in milliseconds. Separate values with commas, spaces, or new lines. Results update as you type.
Comparing 100 baseline and 100 current measurements. Mean: 195 → 195 ms. p95: 195 → 100 ms. p99: 195 → 2,000 ms.
Sorted request times · bar height uses the same scale in both windows
| Measure | Baseline | Current |
|---|---|---|
| Count | 100 | 100 |
| Mean | 195 ms | 195 ms |
| Median (p50) | 195 ms | 100 ms |
| p95 | 195 ms | 100 ms |
| p99 | 195 ms | 2,000 ms |
| Population variance | 0 ms² | 171,475 ms² |
| Standard deviation | 0 ms | 414.1 ms |
| Maximum | 195 ms | 2,000 ms |
A mean can stay flat while the tail changes. Read the count and distribution with each summary; this small sample calculator does not establish an SLO or explain why a request was slow.
Compare like with like, then investigate the change.
Compare the same request population, units, time window, and percentile convention. Relative
change is (current − baseline) / baseline × 100. A release policy can use a stated
threshold to flag a change for investigation—for example, a p95 increase greater than 15%.
A flag is a decision rule, not proof that the release caused the change. Check traffic mix, sample size, errors, and other changes in the same window. Use the threshold to decide when a person should look closer, not as a substitute for understanding what users experienced.
Calculate p95 and use it in a release check.
Work through nearest-rank p95, a fuller latency summary, and a threshold comparison. Each ticket uses the same percentile rule in TypeScript and Go.
Math in Practice practice 8 min
Find the nearest-rank p95
This is an experiment with ticket-style exercises, giving beginners a feel for how tasks may be described in the workplace. Leave feedback
Checking your sign-in status. Your lesson remains available while we check.