01 / The idea
Retrying a blip is a fair start.
You’re building a webhook forwarder: when something happens in your app, it posts an event to a customer’s URL. Networks blip, so when a post fails with a 503 you try again straight away, up to eight times. For one event and a brief failure, that loop recovers before anyone notices.
Read the first retry loopTypeScript · the version this lesson starts from
// The first version: one event, a brief blip, so try again straight away.
export async function forwardWithRetry(send: () => Promise<number>, attempts = 8): Promise<number> {
let status = 0;
for (let attempt = 1; attempt <= attempts; attempt++) {
status = await send();
// Stop on anything that isn't overload or an outage.
if (status < 500 && status !== 429) return status;
}
return status;
} Go’s version is the same loop. Both languages meet again at retryDelay in section
02.
Then the destination goes down while twelve events are in flight. Every event follows the same helpful rule against the same failing service: fail, try again, fail again. Twelve loops that each look careful add up to a steady stream of requests at a service that’s trying to come back. None of them stops because retrying is pointless, only because it ran out of tries.
Backoff means waiting longer after each failure. Jitter means picking a random moment inside that wait, so callers that failed together don’t come back together. And a retry is only safe when the receiver can tell a repeat from new work: that’s idempotency. Marc Brooker’s post on the AWS Architecture Blog is blunt about the second part: jittered backoff “should be considered a standard approach for remote clients.”
If you’ve used TanStack Query, you’ve already shipped this. A failed query retries three
times in the browser, waiting Math.min(1000 * 2 ** failureCount, 30000) milliseconds: one second, two, four, and never more than thirty. Section 05 writes that default
out, then builds a reconnect you have to own.
02 / See the shape
Wait longer, spread out, and know when to stop.
The basic form is the spacing rule: how long to wait before the next try. In the wild adds the decision to try at all, with a reason for every stop. At the call site runs one outage through all three policies.
Both languages run the same model on a logical clock, with no real network or sleeps, and replay the same 18 shared schedules.
The spacing rule. Immediate retries wait one tick, backoff doubles up to a cap, and jitter picks a random point inside that window.
// When to try again: nothing between tries, a doubling wait, or a random point inside it.
export function retryDelay(
policy: Policy,
failedAttempts: number,
random: number
): { delay: number; random: number } {
const window = Math.min(8, 2 ** Math.min(failedAttempts, 3)); // 2, 4, 8, 8… ticks.
if (policy === 'immediate') return { delay: 1, random }; // Next tick: no deliberate backoff.
if (policy === 'backoff') return { delay: window, random };
const next = (Math.imul(random, 1664525) + 1013904223) >>> 0;
const unit = next / 4294967296;
return { delay: Math.max(1, Math.ceil(unit * window)), random: next };
// Full jitter rounded up to our one-tick clock. The first tick is the minimum.
} // When to try again: nothing between tries, a doubling wait, or a random point inside it.
func retryDelay(policy string, failedAttempts int, random uint32) (int, uint32) {
window := 1 << min(failedAttempts, 3) // 2, 4, 8, 8… ticks.
if policy == "immediate" {
return 1, random
} // Next tick: no deliberate backoff.
if policy == "backoff" {
return window, random
}
next := random*1664525 + 1013904223 // uint32 wrapping is intentional.
unit := float64(next) / 4294967296.0
return max(1, int(math.Ceil(unit*float64(window)))), next
// Full jitter rounded up to the model's one-tick clock.
} Reading the TypeScriptA repeatable random stream
retryDelay returns the wait and the next random state, so every event
carries its own repeatable stream. Math.imul multiplies as 32-bit integers,
and >>> 0 turns the result back into an unsigned number. The seed
keeps the film and the tests repeatable; production code can use Math.random().
Unless it stops the event, afterFailure changes only when it’s next due and its
random state. The key and the deadline stay as they were, so a retry is the same operation
again.
Reading the Gouint32 wraps on purpose
random*1664525 + 1013904223 overflows a uint32, and the Go spec defines unsigned
arithmetic as wrapping, which matches TypeScript’s >>> 0.
1 << min(failedAttempts, 3) gives windows of 2, 4, 8, then 8 ticks. min and max have been built in since Go 1.21.
03 / Follow the outage
Watch twelve events meet one outage.
Five steps, every bar from running the forwarder you just read. Each column is one tick: first attempts at the bottom, retries stacked on top, and a strip underneath for whether the destination was up. Before each step, guess how many attempts it takes to deliver all twelve.
In Try it, break and restore the destination yourself, then switch policies on the same history.
One outage, three retry policies.
Retry immediately1 event
One event. Down for 1 tick, then healthy
- attempts
- 2
- delivered
- 1
- stopped
- 0
- most retries in a tick
- 1
Retry immediately, 1 event. One event. Down for 1 tick, then healthy. Attempts per tick: t0: 1, t1: 1. 2 attempts, 1 delivered, 0 stopped, peak 1 retries in a tick. One failure, one retry, delivered. For a single event and a brief blip, retrying straight away is fine.
A blip, one retry.
One event hits a 503, tries again next tick, and gets through.
Reduced motion: choose a scene to see its completed state.
Read this scene
One event hits a 503, tries again next tick, and gets through.
Retry immediately, 1 event. One event. Down for 1 tick, then healthy. Attempts per tick: t0: 1, t1: 1. 2 attempts, 1 delivered, 0 stopped, peak 1 retries in a tick. One failure, one retry, delivered. For a single event and a brief blip, retrying straight away is fine.
Watch restarts when you return. Step through keeps your selected step. Try it starts a fresh batch of twelve events each time you open it.
What backoff and jitter buy you
Now put names on what you just watched. These are the words you’ll hear in a design review, and each one points at something on this page.
- Fewer wasted attempts
- The same outage costs 60 attempts retrying immediately, 48 with backoff, and 39 with jitter. All twelve events are delivered every time.
- Retries that don’t arrive together
- Without jitter, all twelve retry at t2 and again at t6. With it, no tick sees more than 9 retries.
- Bounded work
- An attempt limit and a deadline end every event. In step 5 the destination never comes back, and the forwarder still stops.
- Stops that explain themselves
- Each stop names its reason: a request that can’t change, a repeat that isn’t safe, the attempt limit, or the deadline.
- Safe to repeat, or not repeated
- Every try carries the event’s same key, and the forwarder won’t retry at all unless the receiver can recognize a repeat.
The review words are exponential backoff, jitter, and the thundering herd they break up. The promise that a repeat is harmless is idempotency. Section 08 covers what they cost.
04 / Try a decision
Two careful retry loops multiply.
A worker forwards each event with its own loop of up to three attempts. Later the team
adopts an HTTP client that retries 503s by itself, also up to three attempts. Both loops are
in layered.ts, and the lesson’s tests count what they send.
05 / Give it a real job
The worker owns the retries. The receiver makes them safe.
In a real forwarder, events wait in a durable queue. A worker takes one, sends it, and decides from the reply whether to retry, when, and when to give up. The customer’s endpoint does the other half: it recognizes a repeat by the event’s ID and doesn’t do the work twice.
Owns the budget
One place decides the wait, the attempt limit, and the deadline.
Says when to wait
A 429 or 503 can carry Retry-After, and that wait comes first.
Makes repeats safe
It records the event IDs it has handled and skips a repeat.
An event that stops still needs somewhere to go. A real forwarder records it as failed, with its reason, so someone can fix the endpoint and send it again with the same key. Retry-After can be a number of seconds or a date, and the forwarder treats it as the shortest wait it’s allowed.
The example leaves out the queue, real HTTP, parsing Retry-After, and storing failed events. None of those change who owns the decision to try again.
Build UIs?The data-fetching library you use already retries for you, and one day a live connection makes you write the backoff yourself.
Where it already is in your components
TanStack Query retries a failed query three times in the browser, doubling the wait from one second up to a thirty-second cap, and pauses those retries while the device is offline. You rarely notice, except as a spinner that lasts a few seconds longer before the error.
The textbook panes write those defaults out with one change: they only retry failures that
might pass next time. A 404 won’t, so it shows straight away. Put a retry loop inside queryFn as well and you’re back in section 04, with the attempts multiplying.
When you have to own it
Now it’s a live notifications badge over a WebSocket. When the server restarts, every open tab loses its connection in the same second. If they all reconnect one second later, they arrive together: the crowd from step 3, made of your users’ tabs. So each tab draws a random wait inside a doubling window, capped at thirty seconds.
A connection that opens resets the count. After eight failures in a row the badge stops
and offers Try again instead of retrying forever. While the browser reports it’s offline,
waiting can’t help, so the badge waits for the online event instead. MDN
calls navigator.onLine “inherently unreliable”, so the badge uses it only to pause, never to give up.
// One place decides how a live connection comes back. Components only call these.
export const maxFailures = 8;
// Full jitter: a random wait between zero and a doubling window, capped at 30 seconds.
// Every open tab draws its own wait, so they don't all reconnect in the same second.
export function reconnectDelay(failures: number, random: () => number = Math.random): number {
const window = Math.min(30_000, 1_000 * 2 ** failures);
return Math.round(random() * window);
}
export type NextStep = 'wait-for-network' | 'retry' | 'give-up';
// Waiting can't fix a missing network, and trying forever isn't a plan either.
export function nextStep(failures: number, online: boolean): NextStep {
if (!online) return 'wait-for-network';
return failures >= maxFailures ? 'give-up' : 'retry';
}
An order list whose query retries three times with doubling waits capped at thirty seconds, and only for failures that might pass next time. TanStack Query in React and Svelte.
import { useQuery } from '@tanstack/react-query';
type Order = { id: string; total: number };
class HttpError extends Error {
status: number;
constructor(status: number) {
super(`HTTP ${status}`);
this.status = status;
}
}
export function OrderList() {
const orders = useQuery({
queryKey: ['orders'],
queryFn: async (): Promise<Order[]> => {
const response = await fetch('/api/orders');
if (!response.ok) throw new HttpError(response.status);
return response.json();
},
// Retry what might pass next time: no response, overload, or an outage. Not a 404.
retry: (failureCount, error) => {
const mightChange =
!(error instanceof HttpError) || error.status === 429 || error.status >= 500;
return mightChange && failureCount < 3;
},
// TanStack Query's default wait, written out: 1s, 2s, 4s… capped at 30s.
retryDelay: (failureCount) => Math.min(1000 * 2 ** failureCount, 30_000)
});
if (orders.isPending) return <p>Loading orders…</p>;
if (orders.isError) return <p role="alert">Couldn’t load orders: {orders.error.message}</p>;
return (
<ul>
{orders.data.map((order) => (
<li key={order.id}>{order.total}</li>
))}
</ul>
);
}
06 / Recognize it elsewhere
Anywhere something tries again on your behalf.
You’ve met all of these. For each one, find who decides the wait and what makes a repeat safe.
| Where you’ve seen it | What tries again | How it waits, and what keeps it safe |
|---|---|---|
| TanStack Query | A failed query | Three retries in the browser, doubling from 1s to a 30s cap. Reads are safe to repeat. |
| Stripe webhooks | An event your endpoint didn’t accept | Exponential backoff for up to three days in live mode. Your endpoint skips event IDs it has already handled. |
| Stripe’s API | A POST whose reply never came | You retry with the same Idempotency-Key, and Stripe replays the first
result. |
| A chat or notifications socket | A dropped connection | Your code: a jittered wait that resets once connected. |
One caller retrying once after a blip doesn’t need a policy. It becomes one when many callers can fail at the same moment.
07 / Already in your toolbox
Your tools already retry this way.
Three places to look. For each one, find the wait, the limit, and what makes a repeat safe.
AWS Architecture Blog · Exponential Backoff and Jitter
Marc Brooker’s 2015 simulation behind “full jitter”: sleep = random(0, min(cap, base * 2 ** attempt)). Its numbers come from a
different workload, but the clusters it shows are step 3 of this lesson.
TanStack Query · retryer.ts
The retry loop behind every query: a default of three retries in the browser and none on the server, the doubling delay capped at 30 seconds, and a pause while offline.
Read the retryer source ↗Stripe · Idempotent requests
Stripe saves the status and body of the first request for a key, “regardless of whether it succeeds or fails”, and returns it for retries. Keys can be removed after 24 hours, and reusing one with different parameters is an error.
Read the reference ↗A useful counterexample: a Pay buttonWhen not to retry automatically
A customer taps Pay and the request times out. The charge may already have gone through. Retrying without a key could charge them twice, and a spinner that quietly retries hides the question. Show that the outcome is unknown and check the order first.
With an idempotency key created once for that payment, a retry with the same key is safe. That’s exactly the job Stripe’s keys do.
08 / The parts to watch
Every retry is more work somewhere.
These are the places it still goes wrong.
Retries stack across layers
A client library, an SDK, a proxy, and your own loop may each retry. Their limits multiply, as section 04 showed. Decide which layer owns retries and give the others a single attempt.
A timeout doesn’t mean it didn’t happen
The request may have reached the server and done its work before the reply was lost. Only
retry what the receiver can recognize as a repeat. The forwarder refuses to retry when repeatSafe is false.
Some failures won’t change
A payload the destination rejects as invalid will be rejected again. The forwarder stops those after one attempt. Which replies are final is the API’s contract, not a rule that every 4xx is.
The deadline doesn’t restart
The deadline is fixed when the event is created. Resetting it after each failure would let an event retry forever, one fresh deadline at a time.
Per-event limits don’t cap the total
Eight attempts per event is still eight thousand attempts for a thousand events. A retry budget caps retries across every request to a destination, and a circuit breaker stops sending for a while when failures keep coming. This model has neither.
Random waits make tests flaky
Pass the randomness in, like the forwarder’s seeded stream, so a test can replay one exact schedule.
09 / Make the call
What would you have to change tomorrow?
Give both designs a plausible change and follow the work it creates.
| The change | Retry immediately | Backoff, jitter, and limits |
|---|---|---|
| One request, one brief blip | Retries next tick and gets through first. | Also recovers, after a random wait of one or two ticks. |
| The destination goes down with twelve events in flight | 60 attempts, with all twelve retrying in the same tick. | 39 attempts, and never more than 9 retries in a tick. |
| It stays down | Eight attempts back to back, then an error with no reason. | Each event stops by its deadline and records why. |
| The receiver can’t recognize repeats | Retries anyway, so a lost reply can deliver an event twice. | Refuses to repeat; the event stops as unsafe. |
| Someone adds a retrying HTTP client | Attempts multiply. | Attempts multiply here too. Only choosing one owner fixes it. |
Reach for backoff with jitter when many callers can fail at the same moment, and put a limit and a deadline on every retry. Twelve events meeting one outage is the moment.
Keep the simple loop for one caller retrying one cheap, safe request. A command-line tool fetching a file once doesn’t make a crowd.
The question I’d leave beside the code is: who else is retrying this right now, and is it safe to send twice?
10 / Take the idea with you
Explain the forwarder without saying “backoff.”
“When a send fails, wait before trying again, a bit longer each time and at a random moment so everyone doesn’t come back together. Stop when it can’t work or time’s up, and only repeat what the receiver can recognize as a repeat.” In a review, the words are exponential backoff, jitter, retry budget, and idempotency key.
Before moving on, jot down why twelve careful loops flooded the destination, why two retry layers sent nine requests, and one place in your own code that retries, along with whoever else might be retrying the same request.
Connections to follow nextRelated lessons
- Idempotency and at-least-once delivery is how a receiver makes a repeat harmless, which this lesson assumed.
- Backpressure and queues covers work arriving faster than a destination accepts it. Retries are part of that work.
- Race conditions in UI follows a retry that finishes late and mustn’t overwrite newer data.
- Discriminated unions model the badge’s connected, waiting, offline, and stopped states.
- Timeouts, deadlines, and races decide when a single attempt has taken too long.