← Math in Practice
Concept Math in engineering decisions

Retry amplification and backoff

A retry can help one request and add work to an already failing service.

A report service begins timing out. The API’s incoming rate is steady at 700 requests per second, but the dependency now records far more attempts than user requests. A retry policy may be turning each failure into extra work exactly when the service has the least room for it.

The judgment to keep

Count initial requests separately from attempts sent to a dependency. State the assumptions behind any expected value, and treat retries as added demand with a deadline and a budget.

TypeScriptGo Bounded retries · expected attempts · exponential backoff
01 / Read the incident

A stable caller rate can hide a growing attempt rate.

At 14:10, the API receives about 700 report requests per second. The database-facing client records 1,280 attempts per second, and timeouts are rising. That difference is evidence of extra calls somewhere, but it does not yet prove that retries are the cause. A traffic mix change, duplicate work, queue replay, or a metric boundary mismatch could also explain it.

First confirm that the counters cover the same service path and time window. Count unique incoming request IDs separately from dependency attempt IDs, then inspect how each attempt relates to its parent request.

Case file / Report generationCallers are steady; dependency work has climbed.
Incoming rate
700 logical report requests per second.
Dependency rate
1,280 recorded attempts per second.
Observed outcome
Timeouts and queue depth are rising.
Question
Are retries adding work, and what could the current policy send at peak?
Leave with: a symptom, a competing explanation, and the measurement that could distinguish retry traffic from a boundary or workload change.
02 / Separate the causes

Count parent requests and attempts before changing the policy.

Hold the caller workload and time window fixed. For a sample of parent request IDs, count dependency attempts, note their status and timing, and inspect whether the client created them after retryable failures. Compare with one controlled request under a disposable test dependency that times out on demand. Do not create load against a production dependency to answer this question.

If each parent has one attempt, investigate duplicate upstream work or mismatched counters. If some parents have two or three attempts following failures, retries explain at least part of the excess. If the excess begins before client retries, inspect queues, fan-out, and request mix too.

01 / CallerOne logical request

Counted once at the API boundary.

02 / First callAttempt 1

Sent to the dependency.

03 / FailureRetry decision

Only selected failures may trigger another attempt.

04 / Added workAttempt 2…m

Correlate each attempt with its parent.

Checkpoint: what observation would make you stop blaming retries and look for a different source of extra dependency work?
03 / Count retry work

Retries raise the ceiling and the expected work.

Let p be the probability that one attempt fails, and let m be the maximum number of attempts including the first call. Under the simplifying assumption that each attempt fails independently with the same probability, the expected attempts per logical request are:

E[A] = 1 + p + p² + … + p^(m−1)

Each term is the probability that the corresponding attempt is reached. With p = 0.5 and m = 3, E[A] = 1 + 0.5 + 0.25 = 1.75. At 700 incoming requests per second, that model predicts 700 × 1.75 = 1,225 dependency attempts per second. The hard maximum is 700 × 3 = 2,100 attempts per second if every logical request uses all attempts.

With those assumptions, the probability all three attempts fail is p³ = 12.5%, so the final success probability is 87.5%. Retries buy a chance of recovery by spending more calls and more time. Whether that trade is good depends on the operation, deadline, and dependency state.

Leave with: a best estimate under stated assumptions and a maximum request multiplier for the configured attempt limit.
04 / Spread the attempts

Backoff changes when work arrives, not how much work is allowed.

Immediate retries can synchronize many callers into a burst. Exponential backoff spaces retry number k by dₖ = min(cap, base × 2^(k−1)). With a 100 ms base, 800 ms cap, and four retries, the nominal waits are 100, 200, 400, and 800 ms. Jitter varies the wait so clients do not all retry together; it does not reduce the attempt count by itself.

Backoff also spends the caller’s deadline. If the remaining request budget cannot contain the wait and another attempt, stop retrying. A retry policy should bound attempts, total elapsed time, and the kinds of errors worth retrying.

Illustrative schedule · 100 ms base · 800 ms cap
Retry numberNominal waitWhat doubles
1100 msBase
2200 ms100 × 2
3400 ms100 × 4
4800 msCap reached
Checkpoint: identify the deadline, retryable failures, idempotency guarantee, and attempt budget before selecting a backoff schedule.
05 / Change the assumptions

See expected dependency work and its hard ceiling.

Retry load labLogical requests → dependency attempts

Adjust one assumption at a time. Retries are additional attempts after the first.

Expected attempts per request1.751 + p + p² … through attempt 3
Expected attempt rate1225 attempts/sUnder independent, constant failure probability
Maximum attempt rate2100 attempts/s700 requests/s × 3 attempts

This model excludes correlated failures, timeouts, queueing, retry budgets, and cancellation. Actual incident traffic needs attempt-level measurements.

Try a 100% failure probability. Expected attempts should equal the maximum because every retry is reached. Then try 0%: every request stops after its first attempt. At 50%, the estimate lies between those bounds only under the independence model.

Experiment: if an upstream is already overloaded, why might the expected-attempt estimate understate the added load?
06 / Practice in code

Calculate the model and guard its inputs.

The snippets calculate expected attempts with a short sum, so each term remains visible. They also produce a capped exponential schedule. The schedule is nominal: production jitter, deadlines, cancellation, and retryable error policy belong to the calling system and must be chosen explicitly.

Compare the retry calculation in TypeScript and Go.

Both snippets keep the attempt limit and backoff cap explicit.

TypeScriptExpected attempts and capped backoff
retries.ts
export function expectedAttempts(failureProbability: number, maxAttempts: number): number {
	if (!Number.isFinite(failureProbability) || failureProbability < 0 || failureProbability > 1) {
		throw new Error('failureProbability must be between 0 and 1');
	}
	if (!Number.isInteger(maxAttempts) || maxAttempts < 1) {
		throw new Error('maxAttempts must be a positive integer');
	}
	let probabilityAllPreviousFailed = 1;
	let attempts = 0;
	for (let attempt = 1; attempt <= maxAttempts; attempt += 1) {
		attempts += probabilityAllPreviousFailed;
		probabilityAllPreviousFailed *= failureProbability;
	}
	return attempts;
}

export function backoffDelays(baseMs: number, capMs: number, retries: number): number[] {
	if (
		!Number.isFinite(baseMs) ||
		!Number.isFinite(capMs) ||
		baseMs <= 0 ||
		capMs <= 0 ||
		!Number.isInteger(retries) ||
		retries < 0
	) {
		throw new Error('base and cap must be positive; retries must be a non-negative integer');
	}
	return Array.from({ length: retries }, (_, index) => Math.min(capMs, baseMs * 2 ** index));
}
GoExpected attempts and capped backoff
retries.go
package mathpractice

import (
	"errors"
	"math"
	"time"
)

func ExpectedAttempts(failureProbability float64, maxAttempts int) (float64, error) {
	if math.IsNaN(failureProbability) || failureProbability < 0 || failureProbability > 1 {
		return 0, errors.New("failure probability must be between 0 and 1")
	}
	if maxAttempts < 1 {
		return 0, errors.New("max attempts must be positive")
	}
	probabilityAllPreviousFailed := 1.0
	attempts := 0.0
	for attempt := 1; attempt <= maxAttempts; attempt++ {
		attempts += probabilityAllPreviousFailed
		probabilityAllPreviousFailed *= failureProbability
	}
	return attempts, nil
}

func BackoffDelays(base, cap time.Duration, retries int) ([]time.Duration, error) {
	if base <= 0 || cap <= 0 || retries < 0 {
		return nil, errors.New("base and cap must be positive; retries must be non-negative")
	}
	delays := make([]time.Duration, retries)
	delay := base
	for i := range delays {
		if delay > cap {
			delay = cap
		}
		delays[i] = delay
		if delay <= cap/2 {
			delay *= 2
		} else {
			delay = cap
		}
	}
	return delays, nil
}
07 / Make the next call

Reduce the work before asking a failing service to do more.

A useful incident answer states what is measured, which model assumptions apply, the maximum extra work allowed, and what signal would make the team stop retrying. That is a stronger decision than “the average failure rate is only 10%.”

For practical policy guidance, see the AWS Well-Architected guidance on limiting retries and the Amazon Builders’ Library article on timeouts, retries, and jitter. The HTTP semantics specification’s idempotency section explains when repeating a request has the same intended effect. Checked 2026-09-30; these references inform policy choices but do not establish a particular application’s retry behavior.

Question to keep: how many logical operations can become dependency attempts, how quickly can they retry, and when does the deadline say to stop?