“One in a thousand” is a per-request statement.
During a checkout review, a dependency test reports that 1 out of 1,000 requests timed
out. The team is deciding whether to ship a change that makes the dependency part of every
checkout. The reported fraction is 1 ÷ 1,000 = 0.001 = 0.1% per tested request.
That is an observation from a test window, not automatically the true future failure
probability. Before interpreting it, check how many requests were tested, whether they
match production traffic, and whether failures clustered in time or by region, account, or
request type. For now, use p = 0.001 as an illustrative model input.
- Event
- One checkout request times out at the dependency boundary.
- Illustrative p
- 0.1% = 0.001 per request.
- Window
- 1,000 checkout requests in one busy minute.
- Question
- What count should we expect, and how likely is at least one timeout?
Probability only means something after you say what can happen.
For a single request, let X be 1 when that request times out and 0 when it
does not. This is a Bernoulli trial: a two-outcome model with probability p of the event of interest. Here the illustrative value is p = 0.001.
A probability is a number from 0 to 1 that expresses how often an event would occur over
repeated comparable opportunities under a model. 0.001 is the same fraction
as 0.1%. It does not mean the failure happens on a schedule every thousandth
request, and it does not identify why a timeout occurred.
This example counts dependency timeouts, not failed checkouts. The API might retry, use a fallback, or fail the whole operation; those are different downstream events with different probabilities and consequences. Keep the boundary fixed while doing this calculation.
Expected value is a long-run average, not a promised count.
Let Xᵢ count whether request i times out. Across n requests, the total number of timeouts is S = X₁ + X₂ + … + Xₙ. If every
request has the same timeout probability p, the expected count is E[S] = n × p.
For 1,000 requests at p = 0.001: E[S] = 1,000 × 0.001 = 1 expected timeout. Across many comparable 1,000-request windows, the average count would tend
to one if the rate stayed as modeled. A particular window could have zero, one, or several.
The expected count is a weighted average of outcomes. If 99.9% of requests do not time out
and 0.1% do, the per-request expected timeout indicator is 0 × 0.999 + 1 × 0.001 = 0.001. Add that expectation over 1,000 requests to get one. This linearity does not require
independence; it does require that the per-request probabilities used in the sum apply to
the requests being counted.
| Requests, n | Calculation, n × p | Expected timeouts | Interpretation |
|---|---|---|---|
| 100 | 100 × 0.001 | 0.1 | A fractional expectation is valid; an actual count is whole. |
| 1,000 | 1,000 × 0.001 | 1 | One on average across comparable windows. |
| 10,000 | 10,000 × 0.001 | 10 | Ten on average if the same per-request rate applies. |
“At least one” is a different question from “how many on average?”
For a particular request, the probability of no timeout is 1 − p = 0.999. If
requests fail independently and all share that probability, the chance that none of 1,000
requests time out is 0.999¹⁰⁰⁰ ≈ 0.3677, or about 36.8%.
“At least one” is the complement of “none”: P(S ≥ 1) = 1 − P(S = 0) = 1 − (1 − p)ⁿ. So the modeled probability of at least one timeout is 1 − 0.999¹⁰⁰⁰ ≈ 0.6323, about 63.2%. The expected count is still
one. These are related, but they answer different questions.
The code uses logarithms to evaluate the complement accurately when p is very
small: −expm1(n × log1p(−p)) is algebraically the same expression as 1 − (1 − p)ⁿ, with less loss of precision from subtracting nearby values.
Illustrative timeout probability.
Assumes independent requests with the same p.
Modeled chance of one or more timeouts.
Expected count, not the most likely exact count.
p = 0, the chance of any failure is zero. If p = 1 and n > 0, it is one. If n = 0, no event was
tested, so the chance is zero.The same expected count can hide very different risks.
Expected count uses linearity: if each of 1,000 requests has marginal timeout probability
0.001, the expected total is one even when outcomes are dependent. But the formula for “at
least one,” 1 − (1 − p)ⁿ, assumes independence as well as a shared
probability.
Consider two hypothetical systems with the same expected count. In one, each request fails independently at probability 0.001: a 1,000-request window has about a 63.2% chance of at least one timeout. In another, all 1,000 requests succeed together 99.9% of the time, and all 1,000 time out together 0.1% of the time. The expected count is still one, but the chance of any timeout is only 0.1%; when it happens, the impact is a burst of 1,000 failures.
This perfectly correlated example is deliberately stark. Real systems can sit between these cases: a regional network fault or a dependency restart can make many failures arrive together. Check time series, request cohorts, region, dependency instance, and release version. A single average rate cannot reveal clustering.
Keep the assumptions visible in the function.
Both examples calculate an expected count and the independent-model probability of one or more failures. They reject an invalid probability and a negative or non-integer event count, then handle the boundary values explicitly. The return values are model outputs; they do not infer whether the input estimate is representative.
Both versions return an expected count and use the independent-trial assumption for the window probability.
export type FailureEstimate = {
expectedFailures: number;
probabilityAtLeastOne: number;
};
/**
* Models n independent events that each fail with probability p.
* This is a teaching estimate, not a service-level prediction.
*/
export function estimateFailures(eventCount: number, failureProbability: number): FailureEstimate {
if (!Number.isSafeInteger(eventCount) || eventCount < 0) {
throw new Error('eventCount must be a non-negative safe integer');
}
if (!Number.isFinite(failureProbability) || failureProbability < 0 || failureProbability > 1) {
throw new Error('failureProbability must be between 0 and 1');
}
const expectedFailures = eventCount * failureProbability;
const probabilityAtLeastOne =
eventCount === 0 || failureProbability === 0
? 0
: failureProbability === 1
? 1
: -Math.expm1(eventCount * Math.log1p(-failureProbability));
return { expectedFailures, probabilityAtLeastOne };
}
package mathpractice
import (
"errors"
"math"
)
type FailureEstimate struct {
ExpectedFailures float64
ProbabilityAtLeastOne float64
}
// EstimateFailures models n independent events that each fail with probability p.
// This is a teaching estimate, not a service-level prediction.
func EstimateFailures(eventCount int64, failureProbability float64) (FailureEstimate, error) {
if eventCount < 0 {
return FailureEstimate{}, errors.New("event count must be non-negative")
}
if math.IsNaN(failureProbability) || failureProbability < 0 || failureProbability > 1 {
return FailureEstimate{}, errors.New("failure probability must be between 0 and 1")
}
estimate := FailureEstimate{ExpectedFailures: float64(eventCount) * failureProbability}
switch {
case eventCount == 0 || failureProbability == 0:
estimate.ProbabilityAtLeastOne = 0
case failureProbability == 1:
estimate.ProbabilityAtLeastOne = 1
default:
estimate.ProbabilityAtLeastOne = -math.Expm1(float64(eventCount) * math.Log1p(-failureProbability))
}
return estimate, nil
}
Use the estimate to choose the evidence you need next.
The calculation says that a modeled 0.1% per-request timeout rate is not equivalent to “no one will notice.” At 1,000 requests, the independent model puts the chance of one or more at about 63.2%. That number is useful as a prompt to examine exposure; it is not proof that the rollout will produce a timeout or that timeouts will be independent.
Before deciding, verify whether the test matches production request mix and regions. Then compare candidate causes—new dependency behavior, one unhealthy region, or a shared timeout configuration—against evidence that separates them. During rollout, watch actual timeout counts and rates beside total requests, segment the signal by cohort, and define a stop condition before increasing exposure.