The feed waits for the slowest required shard.
In this illustrative release review, the feed service issues 100 reads at once and combines their results. The shard latency dashboard says 99% of reads are at or below 200 ms. The request trace shows that many feed responses exceed 200 ms. Both can be true: the page waits for the maximum latency among its required parallel reads.
The inputs are teaching assumptions, not production measurements. We will temporarily assume all 100 reads have the same latency distribution and behave independently. Later we will check why real systems often break those assumptions.
- Fan-out
- 100 parallel shard reads per feed request.
- Shard threshold
- 99% finish within 200 ms, according to the component distribution.
- Join rule
- The response waits for every required shard.
- Question
- How often does the entire fan-out finish within 200 ms?
The maximum is below a threshold only when every call is.
Let F(t) be the probability that one dependency finishes at or before latency t. For N independent calls with that same distribution, the
event “the maximum is at or below t” means all calls finish by t: P(max ≤ t) = F(t)N. This is the distribution of the slowest
parallel call, not the sum of all call times.
At the shard’s p99 threshold, F(t) = 0.99. For 100 parallel shards, 0.99100 ≈ 0.366: only about 36.6% of requests have every shard
finish by that threshold in this model. Equivalently, about 63.4% have at least one shard
exceed it. This is not a prediction that the overall p99 equals a specific shard
percentile; it is one threshold calculation under explicit assumptions.
If independent calls have different distributions, the all-finish probability is ∏ Fi(t). If calls are dependent, the product is not valid without
a joint model. Perfectly correlated calls with the same marginal distribution have the
same threshold probability F(t), while shared congestion can create other
patterns.
Start with the request objective, then derive a model-based component target.
Suppose the product asks for 99% of feed requests to finish within a chosen latency
budget. With 100 independent, identical, required parallel reads and no other work, each
read must meet the budget with probability at least 0.991/100 ≈ 0.9998995, or about 99.98995%, for the model to put 99% of requests inside it. That is a per-read
target near p99.99, not p99.
This inversion is a way to allocate a latency budget, not an SLO decomposition that can be declared true by arithmetic. The aggregator adds its own processing and queueing, network latency matters, dependencies may vary by shard, and timeouts can turn slow work into errors. A per-dependency percentile also needs enough samples to estimate reliably in its tail.
See how fan-out changes the chance of meeting one threshold.
This calculator treats the entered percentage as the probability that one dependency meets the same latency threshold. It computes the probability all calls meet it, the probability at least one misses it, and the per-dependency probability implied by a request target. The model assumes independent, identically distributed calls and a request that waits for all of them.
Under independence: pN.
Complement of every dependency meeting the threshold.
Derived as request target1/N, before reserving other latency budget.
This is a threshold model, not a latency simulator. It omits correlation, different shard distributions, aggregator work, queueing, timeouts, retries, and non-required calls. Compare the estimate with end-to-end traces and request-level measurements.
Use the model to choose evidence, not to guess at the culprit.
A long request does not prove that one unusually slow shard caused it. First compare the request duration with the spans on that request’s critical path. If all shard calls start near together and the response ends when the last required span ends, the maximum model is relevant. If calls start in waves, wait on queues, retry, or execute sequentially, trace that structure before applying the equation.
- Define the request population and latency boundary: start event, end event, and exclusions.
- Record fan-out count, required versus optional work, and whether calls overlap.
- Compare request percentiles with child-span percentiles over the same traffic and window.
- Inspect slow traces for shard identity, queue time, retries, timeouts, region, and release.
- Check whether slow children coincide because of shared network, storage, or load.
- Use enough observations to estimate the requested tail; sparse p99.99 samples are unstable.
Keep the probability assumption visible beside the result.
Both functions validate counts and percentages, calculate the all-meet probability, and invert the formula for a request-level target. The code reports a model result; it cannot verify that calls are independent, identical, parallel, or required.
The function names the distribution assumptions and keeps percent units at its boundary.
export type FanOutEstimate = {
allMeetThreshold: number;
atLeastOneMisses: number;
requiredPerDependency: number;
};
/**
* Estimates the probability that every independent dependency meets a latency
* threshold. The dependency observations are assumed identically distributed.
*/
export function estimateFanOut(
dependencyCount: number,
perDependencyPercent: number,
requestTargetPercent: number
): FanOutEstimate {
if (!Number.isInteger(dependencyCount) || dependencyCount < 1) {
throw new RangeError('dependencyCount must be a positive integer');
}
for (const [name, value] of [
['perDependencyPercent', perDependencyPercent],
['requestTargetPercent', requestTargetPercent]
] as const) {
if (!Number.isFinite(value) || value < 0 || value > 100) {
throw new RangeError(`${name} must be between 0 and 100`);
}
}
const perDependency = perDependencyPercent / 100;
const allMeetThreshold =
perDependency === 0 ? 0 : Math.exp(dependencyCount * Math.log(perDependency));
const target = requestTargetPercent / 100;
const requiredPerDependency = target === 0 ? 0 : Math.exp(Math.log(target) / dependencyCount);
const atLeastOneMisses =
perDependency === 0
? 1
: perDependency === 1
? 0
: -Math.expm1(dependencyCount * Math.log(perDependency));
return {
allMeetThreshold,
atLeastOneMisses,
requiredPerDependency
};
}
const estimate = estimateFanOut(100, 99, 99);
console.log(`All 100 meet threshold: ${(estimate.allMeetThreshold * 100).toFixed(2)}%`);
console.log(`At least one misses: ${(estimate.atLeastOneMisses * 100).toFixed(2)}%`);
console.log(
`Per-dependency target for a 99% request target: ${(estimate.requiredPerDependency * 100).toFixed(5)}%`
);
package main
import (
"fmt"
"math"
)
type FanOutEstimate struct {
AllMeetThreshold float64
AtLeastOneMisses float64
RequiredPerDependency float64
}
// EstimateFanOut assumes independent, identically distributed dependencies.
func EstimateFanOut(dependencyCount int, perDependencyPercent, requestTargetPercent float64) (FanOutEstimate, error) {
if dependencyCount < 1 {
return FanOutEstimate{}, fmt.Errorf("dependencyCount must be a positive integer")
}
for name, value := range map[string]float64{
"perDependencyPercent": perDependencyPercent,
"requestTargetPercent": requestTargetPercent,
} {
if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 || value > 100 {
return FanOutEstimate{}, fmt.Errorf("%s must be between 0 and 100", name)
}
}
perDependency := perDependencyPercent / 100
target := requestTargetPercent / 100
allMeetThreshold := 0.0
if perDependency > 0 {
allMeetThreshold = math.Exp(float64(dependencyCount) * math.Log(perDependency))
}
atLeastOneMisses := 1 - allMeetThreshold
if perDependency > 0 && perDependency < 1 {
// Expm1 keeps the small failure probability accurate when p is near 1.
atLeastOneMisses = -math.Expm1(float64(dependencyCount) * math.Log(perDependency))
}
requiredPerDependency := 0.0
if target > 0 {
requiredPerDependency = math.Exp(math.Log(target) / float64(dependencyCount))
}
return FanOutEstimate{
AllMeetThreshold: allMeetThreshold,
AtLeastOneMisses: atLeastOneMisses,
RequiredPerDependency: requiredPerDependency,
}, nil
}
func main() {
estimate, err := EstimateFanOut(100, 99, 99)
if err != nil {
panic(err)
}
fmt.Printf("All 100 meet threshold: %.2f%%\n", estimate.AllMeetThreshold*100)
fmt.Printf("At least one misses: %.2f%%\n", estimate.AtLeastOneMisses*100)
fmt.Printf("Per-dependency target for a 99%% request target: %.5f%%\n", estimate.RequiredPerDependency*100)
}
For the broader percentile vocabulary, see Mean, median, variance, and percentiles. Google's paper The Tail at Scale discusses how rare component delays become consequential in large-scale interactive systems; the calculation here is a simple independent-call threshold model, not a substitute for measuring a production request path.