← Math in Practice
Concept Latency, probability, and service fan-out

Tail latency under fan-out

A request that waits for every parallel answer inherits the slowest one.

A product feed release looks healthy in each shard dashboard: 99% of shard reads finish within 200 ms. Yet customers loading a feed assembled from 100 shards regularly wait longer than 200 ms. The aggregator does not return when the typical shard finishes. It returns when the last required shard does.

The judgment to keep

Measure the user-facing request and inspect its critical path. Under independence, identical parallel calls make the chance that all finish by a threshold equal to F(t)N; component percentiles do not transfer unchanged to the whole request.

TypeScriptGo Maximum of parallel calls · request latency budgets · critical-path diagnosis
01 / Read the incident

The feed waits for the slowest required shard.

In this illustrative release review, the feed service issues 100 reads at once and combines their results. The shard latency dashboard says 99% of reads are at or below 200 ms. The request trace shows that many feed responses exceed 200 ms. Both can be true: the page waits for the maximum latency among its required parallel reads.

The inputs are teaching assumptions, not production measurements. We will temporarily assume all 100 reads have the same latency distribution and behave independently. Later we will check why real systems often break those assumptions.

Case file / Feed latency review One request fans out, then joins every required result.
Fan-out
100 parallel shard reads per feed request.
Shard threshold
99% finish within 200 ms, according to the component distribution.
Join rule
The response waits for every required shard.
Question
How often does the entire fan-out finish within 200 ms?
02 / Model the fan-out

The maximum is below a threshold only when every call is.

Let F(t) be the probability that one dependency finishes at or before latency t. For N independent calls with that same distribution, the event “the maximum is at or below t” means all calls finish by t: P(max ≤ t) = F(t)N. This is the distribution of the slowest parallel call, not the sum of all call times.

At the shard’s p99 threshold, F(t) = 0.99. For 100 parallel shards, 0.99100 ≈ 0.366: only about 36.6% of requests have every shard finish by that threshold in this model. Equivalently, about 63.4% have at least one shard exceed it. This is not a prediction that the overall p99 equals a specific shard percentile; it is one threshold calculation under explicit assumptions.

If independent calls have different distributions, the all-finish probability is ∏ Fi(t). If calls are dependent, the product is not valid without a joint model. Perfectly correlated calls with the same marginal distribution have the same threshold probability F(t), while shared congestion can create other patterns.

One shard meets 200 msF(t) = 0.99
All 100 meet it, if independent0.99100 ≈ 0.366
At least one exceeds it1 − 0.366 ≈ 63.4%
PopulationFeed requests with 100 required parallel reads
03 / Set a latency budget

Start with the request objective, then derive a model-based component target.

Suppose the product asks for 99% of feed requests to finish within a chosen latency budget. With 100 independent, identical, required parallel reads and no other work, each read must meet the budget with probability at least 0.991/100 ≈ 0.9998995, or about 99.98995%, for the model to put 99% of requests inside it. That is a per-read target near p99.99, not p99.

This inversion is a way to allocate a latency budget, not an SLO decomposition that can be declared true by arithmetic. The aggregator adds its own processing and queueing, network latency matters, dependencies may vary by shard, and timeouts can turn slow work into errors. A per-dependency percentile also needs enough samples to estimate reliably in its tail.

04 / Change the assumptions

See how fan-out changes the chance of meeting one threshold.

This calculator treats the entered percentage as the probability that one dependency meets the same latency threshold. It computes the probability all calls meet it, the probability at least one misses it, and the per-dependency probability implied by a request target. The model assumes independent, identically distributed calls and a request that waits for all of them.

Set the model
All 100 dependencies meet threshold 36.6032%

Under independence: pN.

At least one misses 63.3968%

Complement of every dependency meeting the threshold.

Per-dependency chance needed for 99.00% request target 99.98995%

Derived as request target1/N, before reserving other latency budget.

This is a threshold model, not a latency simulator. It omits correlation, different shard distributions, aggregator work, queueing, timeouts, retries, and non-required calls. Compare the estimate with end-to-end traces and request-level measurements.

05 / Find the real tail

Use the model to choose evidence, not to guess at the culprit.

A long request does not prove that one unusually slow shard caused it. First compare the request duration with the spans on that request’s critical path. If all shard calls start near together and the response ends when the last required span ends, the maximum model is relevant. If calls start in waves, wait on queues, retry, or execute sequentially, trace that structure before applying the equation.

  • Define the request population and latency boundary: start event, end event, and exclusions.
  • Record fan-out count, required versus optional work, and whether calls overlap.
  • Compare request percentiles with child-span percentiles over the same traffic and window.
  • Inspect slow traces for shard identity, queue time, retries, timeouts, region, and release.
  • Check whether slow children coincide because of shared network, storage, or load.
  • Use enough observations to estimate the requested tail; sparse p99.99 samples are unstable.
06 / Practice in code

Keep the probability assumption visible beside the result.

Both functions validate counts and percentages, calculate the all-meet probability, and invert the formula for a request-level target. The code reports a model result; it cannot verify that calls are independent, identical, parallel, or required.

Compare the same model in TypeScript and Go.

The function names the distribution assumptions and keeps percent units at its boundary.

TypeScriptFan-out threshold estimates · independent identical calls
fanout.ts
export type FanOutEstimate = {
	allMeetThreshold: number;
	atLeastOneMisses: number;
	requiredPerDependency: number;
};

/**
 * Estimates the probability that every independent dependency meets a latency
 * threshold. The dependency observations are assumed identically distributed.
 */
export function estimateFanOut(
	dependencyCount: number,
	perDependencyPercent: number,
	requestTargetPercent: number
): FanOutEstimate {
	if (!Number.isInteger(dependencyCount) || dependencyCount < 1) {
		throw new RangeError('dependencyCount must be a positive integer');
	}
	for (const [name, value] of [
		['perDependencyPercent', perDependencyPercent],
		['requestTargetPercent', requestTargetPercent]
	] as const) {
		if (!Number.isFinite(value) || value < 0 || value > 100) {
			throw new RangeError(`${name} must be between 0 and 100`);
		}
	}

	const perDependency = perDependencyPercent / 100;
	const allMeetThreshold =
		perDependency === 0 ? 0 : Math.exp(dependencyCount * Math.log(perDependency));
	const target = requestTargetPercent / 100;
	const requiredPerDependency = target === 0 ? 0 : Math.exp(Math.log(target) / dependencyCount);

	const atLeastOneMisses =
		perDependency === 0
			? 1
			: perDependency === 1
				? 0
				: -Math.expm1(dependencyCount * Math.log(perDependency));

	return {
		allMeetThreshold,
		atLeastOneMisses,
		requiredPerDependency
	};
}

const estimate = estimateFanOut(100, 99, 99);
console.log(`All 100 meet threshold: ${(estimate.allMeetThreshold * 100).toFixed(2)}%`);
console.log(`At least one misses: ${(estimate.atLeastOneMisses * 100).toFixed(2)}%`);
console.log(
	`Per-dependency target for a 99% request target: ${(estimate.requiredPerDependency * 100).toFixed(5)}%`
);
GoFan-out threshold estimates · independent identical calls
fanout.go
package main

import (
	"fmt"
	"math"
)

type FanOutEstimate struct {
	AllMeetThreshold      float64
	AtLeastOneMisses      float64
	RequiredPerDependency float64
}

// EstimateFanOut assumes independent, identically distributed dependencies.
func EstimateFanOut(dependencyCount int, perDependencyPercent, requestTargetPercent float64) (FanOutEstimate, error) {
	if dependencyCount < 1 {
		return FanOutEstimate{}, fmt.Errorf("dependencyCount must be a positive integer")
	}
	for name, value := range map[string]float64{
		"perDependencyPercent": perDependencyPercent,
		"requestTargetPercent": requestTargetPercent,
	} {
		if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 || value > 100 {
			return FanOutEstimate{}, fmt.Errorf("%s must be between 0 and 100", name)
		}
	}

	perDependency := perDependencyPercent / 100
	target := requestTargetPercent / 100
	allMeetThreshold := 0.0
	if perDependency > 0 {
		allMeetThreshold = math.Exp(float64(dependencyCount) * math.Log(perDependency))
	}
	atLeastOneMisses := 1 - allMeetThreshold
	if perDependency > 0 && perDependency < 1 {
		// Expm1 keeps the small failure probability accurate when p is near 1.
		atLeastOneMisses = -math.Expm1(float64(dependencyCount) * math.Log(perDependency))
	}
	requiredPerDependency := 0.0
	if target > 0 {
		requiredPerDependency = math.Exp(math.Log(target) / float64(dependencyCount))
	}

	return FanOutEstimate{
		AllMeetThreshold:      allMeetThreshold,
		AtLeastOneMisses:      atLeastOneMisses,
		RequiredPerDependency: requiredPerDependency,
	}, nil
}

func main() {
	estimate, err := EstimateFanOut(100, 99, 99)
	if err != nil {
		panic(err)
	}
	fmt.Printf("All 100 meet threshold: %.2f%%\n", estimate.AllMeetThreshold*100)
	fmt.Printf("At least one misses: %.2f%%\n", estimate.AtLeastOneMisses*100)
	fmt.Printf("Per-dependency target for a 99%% request target: %.5f%%\n", estimate.RequiredPerDependency*100)
}

For the broader percentile vocabulary, see Mean, median, variance, and percentiles. Google's paper The Tail at Scale discusses how rare component delays become consequential in large-scale interactive systems; the calculation here is a simple independent-call threshold model, not a substitute for measuring a production request path.