← Math in Practice
Concept Change, uncertainty, and evidence

Distributions and long tails

A small share of slow requests can shape what users remember and what capacity must absorb.

The checkout team has a familiar release report: median latency is nearly unchanged, but p99 has climbed. Errors have not moved. Before raising a timeout or adding servers, trace what the measurements say about the requests at the edge of the distribution—and which evidence could explain the change.

The judgment to keep

Treat a distribution as evidence about a population in a particular window. A mean describes average work, a percentile locates a threshold, and neither one diagnoses the cause. Keep counts, cohorts, and the measurement path beside the summary.

TypeScriptGo Skew · latency percentiles · sample size · queue pressure
01 / Read the release signal

The center can stay calm while the slowest users wait longer.

Imagine the same checkout endpoint, route, and region before and after a release. In two illustrative windows of 10,000 eligible requests, the median is 120 ms in both. The first window's p99 is 450 ms; the second's is 1,800 ms. These are teaching values, not production measurements. They describe a change worth investigating; they do not establish that the release caused it.

Start by asking how the requests divide up. Did a new database lookup affect only a route or cohort? Did retries, a cold cache, a downstream slowdown, or a burst create a few long waits? Did trace sampling or aggregation change? “The tail got slower” is an observation to explain, not yet a remedy.

Case file / Checkout release review The typical request and the tail tell different parts of the story.
Population
Eligible checkout requests in matched 10-minute windows.
Sample count
10,000 per window, illustrative and assumed fully retained.
Median
120 ms before and after the release.
p99
450 ms before; 1,800 ms after, using the same percentile convention.
Diagnostic pause: what observation would distinguish a whole-service slowdown from one request path or customer cohort entering the slow tail?
02 / Draw the distribution

Latency is often asymmetric because requests take different paths.

A distribution records how often measurements fall at different values. For latency, values cannot drop below zero, but they can stretch far upward: most requests may finish quickly while a small group waits on a lock, a remote service, disk, garbage collection, a retry, or a queue. Mixing request types with different costs can create the same right-skewed shape even when no single request type has an extreme tail.

A long right tail is an operational description: comparatively few requests take much longer than the bulk. In statistics, heavy-tailed has more specific meanings about how tail probability decays. A handful of slow observations is not enough to prove a power law or any particular distribution family. Do not assume latency is normal, log-normal, or heavy-tailed without evidence and a stated model.

01 / BulkCommon fast path

Cache hits and short queries cluster near the center.

02 / MixtureDifferent work

Large payloads or expensive routes form another cohort.

03 / DelayShared contention

Queues, locks, and downstream waits stretch some calls.

04 / TailRare combinations

Retries or coincident slow dependencies compound a request.

03 / Interpret the tail

A percentile locates an observation; it does not explain the observations above it.

Sort n durations from smallest to largest. Under the nearest-rank convention, the pth percentile is the value at one-based position ceil((p / 100) × n). With 10,000 requests, p99 is at position 9,900: about 99% of observations are at or below that value and the upper one percent are at or above the boundary, subject to ties. Another tool may interpolate between values, so check its convention before comparing results.

The mean answers a different question: total elapsed request time divided by request count. It is useful for aggregate demand and comparisons, but distant values pull it upward. The median (p50) is the midpoint of ordered observations and is less moved by a few extremes. Neither statistic is “the real latency.” A percentile is not the maximum, and p99 is not a guarantee that 99% of future requests will meet that value.

MeanΣ request durations ÷ request count
Nearest-rank p99, n = 10,000ceil(0.99 × 10,000) = observation 9,900
Share above a thresholdcount(duration > threshold) ÷ count
04 / Check what was measured

Before comparing two tails, make sure they describe comparable requests.

Averages and percentiles inherit the choices in the telemetry pipeline. A dashboard might use server processing time while a client measures end-to-end wait. A trace collector might sample one request class more heavily, drop overloaded periods, or merge distributions into sketches. Those choices can shift the visible tail. Record the service boundary, units, time window, request count, inclusion rules, and aggregation method beside the statistic.

Compare like with like first: same endpoint or named cohort, comparable request mix, equal window duration, same percentile implementation, and similar measurement coverage. Then split by region, route, payload class, downstream dependency, and retry outcome. A change in the aggregate can come from a change in request composition, a change inside one cohort, or both.

05 / Connect shape to capacity

Tail latency is a symptom to investigate; average demand still matters for capacity.

A slow request can occupy a worker, connection, or lock longer, keeping resources busy while other requests arrive. If arrivals bunch up or the service is near its sustainable processing limit, a queue can grow and add waiting time to later requests. That feedback makes latency especially sensitive near saturation. But a request's end-to-end latency is not automatically its CPU service time: time spent waiting on a dependency or network may consume a different resource. Measure arrival rate, resource service demand, concurrency, queue depth, and utilization for the resource you are sizing.

The mean helps estimate aggregate work when paired with throughput and resource cost. A tail percentile helps assess whether a slow fraction threatens a user-facing objective. Neither alone gives safe capacity. Use load tests and production evidence across the expected workload, including bursts and dependency behavior, and keep headroom for variability. A p99 crossing a latency objective may justify a page or rollout pause under an agreed policy; it does not by itself say whether the right fix is more capacity, less contention, fewer retries, or isolation of a costly route.

06 / Change the sample

See how the same slow share looks at different sample sizes.

This small discrete model creates a sample where most requests take 100 ms and a chosen share takes longer. Change the sample size and tail share. Watch the mean, p99, and maximum respond. It is a hand-checkable model of order statistics, not a realistic simulation of a service or a prediction about your traffic.

Requests in the tail10 / 1000 (1.00%)
Mean119.0 ms
p99 · rank 990100 ms
Maximum2000 ms

At this sample size the p99 rank still lands among the 990 fast requests. The slow observations are real in this model, but p99 does not include every request above its threshold. The maximum reveals one endpoint, not how often it occurs.

07 / Practice in code

Summarize observations without hiding the conventions.

The TypeScript and Go examples use the same deterministic observations: 990 requests at 100 ms and 10 at 2,000 ms. They report the mean, nearest-rank p99, maximum, sample count, and number at or above a 1,000 ms threshold. The values are synthetic so every result can be checked by hand.

Compare the same summary in TypeScript and Go.

Both examples use nearest rank and preserve the count behind the reported tail.

TypeScriptSummarize a latency sample
distributions.ts
export type LatencySummary = {
	count: number;
	meanMs: number;
	p99Ms: number;
	maxMs: number;
	aboveThreshold: number;
};

/** Nearest-rank quantile summary for one complete, retained sample. */
export function summarizeLatency(observationsMs: number[], thresholdMs: number): LatencySummary {
	if (observationsMs.length === 0) {
		throw new Error('at least one observation is required');
	}
	if (!Number.isFinite(thresholdMs) || thresholdMs < 0) {
		throw new Error('thresholdMs must be finite and non-negative');
	}
	for (const value of observationsMs) {
		if (!Number.isFinite(value) || value < 0) {
			throw new Error('latencies must be finite and non-negative');
		}
	}

	const sorted = [...observationsMs].sort((a, b) => a - b);
	const totalMs = sorted.reduce((sum, value) => sum + value, 0);
	const p99Index = Math.ceil(0.99 * sorted.length) - 1;
	return {
		count: sorted.length,
		meanMs: totalMs / sorted.length,
		p99Ms: sorted[p99Index],
		maxMs: sorted[sorted.length - 1],
		aboveThreshold: sorted.filter((value) => value >= thresholdMs).length
	};
}

const observations = [...Array<number>(990).fill(100), ...Array<number>(10).fill(2000)];
const summary = summarizeLatency(observations, 1000);

console.log(summary);
// { count: 1000, meanMs: 119, p99Ms: 100, maxMs: 2000, aboveThreshold: 10 }
GoSummarize a latency sample
distributions.go
package main

import (
	"errors"
	"fmt"
	"math"
	"sort"
)

type LatencySummary struct {
	Count          int
	MeanMs         float64
	P99Ms          float64
	MaxMs          float64
	AboveThreshold int
}

// SummarizeLatency uses the nearest-rank p99 convention on a complete sample.
func SummarizeLatency(observationsMs []float64, thresholdMs float64) (LatencySummary, error) {
	if len(observationsMs) == 0 {
		return LatencySummary{}, errors.New("at least one observation is required")
	}
	if math.IsNaN(thresholdMs) || math.IsInf(thresholdMs, 0) || thresholdMs < 0 {
		return LatencySummary{}, errors.New("thresholdMs must be finite and non-negative")
	}

	values := append([]float64(nil), observationsMs...)
	var total float64
	aboveThreshold := 0
	for _, value := range values {
		if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
			return LatencySummary{}, errors.New("latencies must be finite and non-negative")
		}
		total += value
		if value >= thresholdMs {
			aboveThreshold++
		}
	}
	sort.Float64s(values)
	p99Index := int(math.Ceil(0.99*float64(len(values)))) - 1

	return LatencySummary{
		Count:          len(values),
		MeanMs:         total / float64(len(values)),
		P99Ms:          values[p99Index],
		MaxMs:          values[len(values)-1],
		AboveThreshold: aboveThreshold,
	}, nil
}

func main() {
	observations := make([]float64, 0, 1000)
	for i := 0; i < 990; i++ {
		observations = append(observations, 100)
	}
	for i := 0; i < 10; i++ {
		observations = append(observations, 2000)
	}

	summary, err := SummarizeLatency(observations, 1000)
	if err != nil {
		panic(err)
	}
	fmt.Printf("count=%d mean=%.1fms p99=%.0fms max=%.0fms above=%.0fms:%d\n",
		summary.Count, summary.MeanMs, summary.P99Ms, summary.MaxMs, 1000.0, summary.AboveThreshold)
	// count=1000 mean=119.0ms p99=100ms max=2000ms above=1000ms:10
}
08 / Choose the next measurement

Make the next query capable of separating plausible causes.

For the checkout example, first verify equivalent windows and retained counts. Then compare latency distributions by route, region, payload, and downstream call. Align those cohorts with queue depth, pool waits, retries, resource utilization, and release changes. If one cohort accounts for the tail, investigate its path. If all cohorts shift together, inspect shared resources and dependencies. If only the telemetry method changed, reproduce the comparison on consistent observations.

A useful incident note separates observation from inference: “In matched 10-minute windows of 10,000 eligible requests, median latency stayed at 120 ms while nearest-rank p99 moved from 450 ms to 1,800 ms. The cause is unknown. We are checking route cohorts, queue waits, and trace coverage before changing concurrency.” That statement is useful because the evidence and the unknown are both explicit.

Transfer: if p99 worsens only for large uploads, but request-weighted p99 stays flat because uploads are rare, which user-facing and capacity measures would you report together? What changes if those uploads become 20% of traffic?