← Architecture
Failure and evidence Before the incident

Thinking in failure modes

Slow, down, wrong, twice, half done.

Every call your code makes to something it does not run can be slow, down, wrong, repeated, or half done. A demo only ever shows the sixth case, where everything answers. Let’s take a gift-card checkout that works, break each of its calls on purpose, and write down what it should do.

The skill to keep: Ask five questions of every call that crosses a boundary, give each answer a policy, and make each one happen in a test before a buyer makes it happen for you.

TypeScriptGoOne checkout, three calls, two recorded builds.

01 / The prompt

“Build the checkout for our gift cards.”

You ask an agent for the online checkout of a café chain’s gift cards. Ask a third-party fraud check whether to go ahead, charge the card through the card processor, create the gift card, show its code. What comes back works. You buy a $50 card, the fraud check answers in a fifth of a second, and the code appears. Every test passes.

Nothing in those first ten minutes tells you what the checkout does on the day the fraud check takes 30 seconds to answer. The default is not reassuring. Go’s HTTP client says it in one line: “A Timeout of zero means no timeout” (net/http, Client), and zero is what you get unless you set it. Node’s built-in fetch comes from undici, which waits up to 300 seconds for a reply’s headers before it gives up (headersTimeout, default 300e3). Both defaults were checked on 23 September 2026.

So the checkout waits as long as the fraud check does. Meanwhile the buyer is looking at a spinner, and the button next to it still says Pay.

The prompt never said how long to wait, what to do without a score, or what a second tap means. The agent had to pick something for each, and a demo cannot show you what it picked.

02 / Name the move

Five questions for every call you do not control.

A failure mode is one specific way a call can go wrong, together with what your system does when it happens. Thinking in failure modes is the habit of listing every call that crosses a boundary and asking the same five questions of each one before the code ships.

What if it is slow? Down? Wrong? Sent twice? Half done? Give every answer a policy, and a test that makes it happen.

The gift-card checkout makes three calls. Here is who runs the other side of each, and what the checkout cannot know about it.

The calls the checkout makes, in order
CallWho runs the other sideWhat the checkout cannot know
Score the purchaseThe fraud vendorWhether it will answer, how soon, and in which shape.
Charge the cardThe card processorWhether a charge happened, when no reply comes back.
Write the cardThe café, in its own databaseNothing, until the process stops between the charge and this write.

Words to put in a prompt or a review

Deadline
How long you will wait for a call before you decide without it.
Fail open, fail closed
Without an answer, go ahead anyway, or stop. Name which, and when.
Idempotency key
A name for one operation, so the other side can recognize a repeat of it.
Outcome unknown
No reply after a side effect. It may have happened. Say so; do not guess.
Ledger
Your record of each step, written before the call it guards.
Failure inventory
The table of calls, conditions, policies, and tests. The thing you hand over.
Is this FMEA?The engineering version, and why this one is smaller

Failure mode and effects analysis comes from reliability engineering. It lists how each component can fail and what each failure does to the whole system, and it usually ranks them. The habit here is the same idea sized for one request path: the rows are the calls your code makes, the five questions are fixed, and the output is a table of policies with a test behind each row. Rank them if you have many; for three calls, write them all.

03 / Follow one checkout

Watch the same bad day hit two checkouts.

First the hopeful build, on the day the fraud check takes 30 seconds. Then the inventoried build on the same day. Then the inventoried build again, when the processor charges the card and the reply never arrives. Step through at your own pace, or open Try it and break the calls yourself.

Failure and evidence

What does a checkout do while a dependency is slow?

Hopeful build

Buyer’s app

Paying $50…

Checkout

This tap has waited
0 ms

Charges 0Cards 0

  • Fraud check

    Not asked yet

  • Card processor

    Not asked yet

  • Gift-card store

    Not asked yet

Buyer → checkoutPOST /checkout {"checkoutId":"gc-1","amountCents":5000}

01/ 03
Thirty seconds, twice

The buyer taps Pay.

The fraud check is having a bad day. The hopeful build asks it with no deadline, because nothing in the prompt mentioned one.

Reduced motion: choose a scene to see its completed state.

Read this scene

The fraud check is having a bad day. The hopeful build asks it with no deadline, because nothing in the prompt mentioned one.

Hopeful build. The buyer’s screen: Paying $50…. This tap has waited 0 ms. Charges at the processor: 0. Cards issued: 0. Fraud check: Not asked yet. Card processor: Not asked yet. Gift-card store: Not asked yet.

On the wire: POST /checkout {"checkoutId":"gc-1","amountCents":5000}

Watch restarts the story when you come back. Step through keeps your step. Try it builds a fresh world for both checkouts every time you run it.

04 / Read it in code

A deadline, a shape check, a key, and a ledger.

Basic form answers three questions for one call: slow, down, wrong. In the wild answers the other two for the whole checkout: twice and half done. At the call site the deadline stops being a number and becomes something the network layer enforces.

Notice what the checkout does not do: it never guesses. No score is its own case with its own policy. No reply from the processor is “pending”, not “failed”, because a failure you report wrongly makes the buyer pay again.

The fraud call with the first three questions answered: a deadline for slow, a status check for down, and a shape check for wrong. Anything that is not a whole-number score from 0 to 100 is not a score.

TypeScriptReading
checkout.ts
export const FRAUD_DEADLINE_MS = 2000;

export type Verdict =
	{ kind: 'scored'; score: number } | { kind: 'timeout' | 'unavailable' | 'invalid' };

/** Ask the fraud check, but only for so long, and believe only a reply that makes sense. */
export async function scoreWithin(deps: Deps, amountCents: number): Promise<Verdict> {
	const reply = await deps.score(amountCents, FRAUD_DEADLINE_MS);
	if (reply.timedOut) return { kind: 'timeout' };
	if (reply.status !== 200) return { kind: 'unavailable' };
	const score = field(reply.body, 'score');
	if (typeof score !== 'number' || !Number.isInteger(score) || score < 0 || score > 100)
		return { kind: 'invalid' };
	return { kind: 'scored', score };
}
GoAlongside
main.go
const FraudDeadline = 2 * time.Second

// Verdict is "scored", "timeout", "unavailable", or "invalid".
type Verdict struct {
	Kind  string
	Score int
}

// ScoreWithin asks the fraud check, but only for so long, and believes only a reply that makes sense.
func ScoreWithin(deps Deps, amountCents int) Verdict {
	reply := deps.Score(amountCents, FraudDeadline)
	if reply.TimedOut {
		return Verdict{Kind: "timeout"}
	}
	if reply.Status != http.StatusOK {
		return Verdict{Kind: "unavailable"}
	}
	var body struct {
		Score *float64 `json:"score"`
	}
	if json.Unmarshal([]byte(reply.Body), &body) != nil || body.Score == nil {
		return Verdict{Kind: "invalid"}
	}
	score := *body.Score
	if score != float64(int(score)) || score < 0 || score > 100 {
		return Verdict{Kind: "invalid"}
	}
	return Verdict{Kind: "scored", Score: int(score)}
}
The behavior these examples promiseChecked by 80 shared cases in TypeScript and Go
  • The fraud call has a 2-second deadline. A reply counts as a score only if it is a 200 whose score is a whole number from 0 to 100. A score of 70 or more declines.
  • Without a usable score, a card of $25 or less is charged; a larger one is held for review with a 202 and nothing is charged.
  • The charge is sent with the checkoutId as its idempotency key and a 5-second deadline. Any reply without a charge id is 202 pending, and the ledger stays at charging.
  • A checkoutId with a final reply gets that reply again, marked as stored, with no calls. One left at charging asks the processor again. One left at charged writes the card.
  • The hopeful build waits without a deadline, reads a missing score as 0, charges with no key, and keeps no ledger.

Every expectation in the shared cases was produced by a separate model written from these rules, kept beside the examples in examples/model/, not copied from either implementation. It covers every combination of five fraud conditions, three processor conditions, a stop before the card, two amounts, and one or two taps.

Reading the TypeScriptunknown, a union, and AbortSignal.timeout

field returns unknown, so the score has to be proved a number before anything compares it. Verdict is a union on kind, so “no score” cannot be mistaken for a score of zero, which is exactly the mistake the hopeful build’s { score = 0 } makes. AbortSignal.timeout rejects the fetch with a TimeoutError DOMException (MDN), which is how the adapter tells a timeout from a refused connection.

Reading the GoA pointer to tell missing from zero, and a context

Decoding into *float64 leaves the pointer nil when the field is missing, so a missing score and a score of 0 stay different. The hopeful build decodes into a plain int and cannot tell them apart. context.WithTimeout puts the deadline on the request, and errors.Is(err, context.DeadlineExceeded) says it was the deadline that ended it. A deadline of 0 means none, the same convention as http.Client.

Run it yourselfNo dependencies

Copy the complete TypeScript file and run node --experimental-strip-types checkout.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:

go.mod
module heyrian.dev/lessons/failure-modes

go 1.23
gc-1 $50, fraud answers: 200 issued GC-0001 after 500 ms | charges 1, cards 1
gc-2 $50, fraud slow: 202 review after 2000 ms, then 202 review (stored) after 0 ms | charges 0, cards 0
gc-3 $20, fraud slow: 200 issued GC-0001 after 2320 ms | charges 1, cards 1
gc-4 $50, charge reply lost: 202 pending after 480 ms, then 200 issued GC-0001 after 320 ms | charges 1, cards 1
gc-5 $50, stopped before the card: 502 stopped after 480 ms, then 200 issued GC-0001 after 20 ms | charges 1, cards 1

The times are the simulated world’s clock, not a measurement: each fake says how long it would take, and the clock adds it up.

05 / Review the agent’s diff

“Checkout no longer hangs when a dependency is slow.”

The complaint was real, and adding timeouts is the right instinct. Read what else changed before you decide. The question to ask of each call is not “can this fail?” but “what happens if it is repeated?”

The agent’s pull request

“Checkout no longer hangs when the fraud check or the processor is slow. Every call now has a timeout and three tries. All 31 tests pass.”

// server/checkout.ts
			(removed)const scored = await fraud.score(amountCents);
			(added)const scored = await withRetry(() => fraud.score(amountCents, { timeoutMs: 2000 }), 3);
			if (scored.score >= 70) return declined();
			(removed)const charge = await processor.charge(amountCents);
			(added)const charge = await withRetry(() => processor.charge(amountCents, { timeoutMs: 5000 }), 3);
			return issueCard(checkoutId, charge.chargeId);
			(added)
			(added)async function withRetry<T>(call: () => Promise<T>, attempts: number) {
			(added)  for (let i = 1; ; i++) {
			(added)    try { return await call(); }
			(added)    catch (error) { if (i === attempts) throw error; }
			(added)  }
			(added)}
			
You are reviewing this change. What do you do?

06 / How it fails

The inventory: every call, every question, every row tested.

This is the artifact the procedure produces. Each row is a call and a condition, what the buyer sees, and what the inventoried checkout does. The last column comes from the shared cases, and the two-tabs row from a native test in each language, so each row is a test that fails if the policy changes.

The gift-card checkout’s failure inventory, for a $50 card
Call and conditionWhat the buyer seesWhat the checkout doesHopeful → inventoried
Fraud checkSlow: takes 30 s“We’re checking this order. No charge yet.” after 2 s. A $20 card goes through.Stops waiting at 2 s. With no usable score, $25 and under goes through and anything more waits for a person.2 charges, 2 cards → 0 charges, 0 cards
Fraud checkDown: refuses the connectionThe same as slow, sooner.The same policy. It cannot tell a dead vendor from a slow one, and does not need to.0 charges, 0 cards → 0 charges, 0 cards
Fraud checkWrong: answers {"risk":"high"}“We’re checking this order.”Believes only a whole-number score from 0 to 100. Anything else counts as no score.1 charge, 1 card → 0 charges, 0 cards
Fraud checkAnswers a high score, 91“This payment was declined.”Declines and stores the answer. Both builds agree here.0 charges, 0 cards → 0 charges, 0 cards
ChargeDown: refuses the connection“Payment pending. Don’t pay again.”Replies 202 pending and keeps the ledger at charging, so the next tap asks the processor again under the same key.0 charges, 0 cards → 0 charges, 0 cards
ChargeHalf done: charged, reply lost“Payment pending,” then the card on the next tap.Retries with the same key. The processor returns the charge it already made.2 charges, 1 card → 1 charge, 1 card
Whole checkoutTwice: the buyer taps Pay againThe same reply both times.Answers from the ledger without calling anything.2 charges, 2 cards → 1 charge, 1 card
Whole checkoutConflicting: two open tabs pay at onceThe same card code in both tabs.Both requests carry the same checkout id, so the charge and the card are made under one key: the processor and the card store return what they already made. Each tab still asks the fraud check once.2 charges, 2 cards → 1 charge, 1 card (native test)
Write the cardHalf done: the process stops after the chargeAn error, then the card on the next tap.The ledger says charged, so the retry skips the charge and writes the card.2 charges, 1 card → 1 charge, 1 card

Five rows share one idea: a side effect may be repeated only if the other side can recognize the repeat. The key does that at the processor, and the ledger does it in the checkout. The rest share another: no answer is its own answer. Decide in advance what it means, instead of letting a default decide.

When a row changes nothing, as the high-score row does here, keep it. It records that you asked. When a row’s test passes on the hopeful build too, check the fake before you celebrate: a fake that never loses a reply cannot tell you what a lost reply does.

Each of these has a lesson of its own: Timeouts, deadlines, and races, Idempotency and at-least-once, and Retry, backoff, and idempotency.

07 / Is it worth it?

You pay for a ledger and a review queue. Here is what they buy.

The hopeful build is shorter and just as fast on a good day. Hold both up against the changes every checkout eventually gets.

The same four changes, made to each build
ChangeHopeful buildInventoried build
A second client: a kiosk in the caféThe kiosk’s own retries charge twice, the same way a second tap does.The kiosk sends a checkoutId and gets the same guarantees.
Replace the fraud vendorA differently shaped reply reads as a score of 0, and every card goes through.A new shape is “no score”: held for review, and visible in the review count.
Change a rule: review above $10, not $25There is no review rule to change.One constant. No other difference.
A payments team takes over the chargeNothing written down about what a retry may do.They inherit the inventory: each row a policy and a test.

Before you change a checkout like this, write down what you will measure and what result you would accept, so the change is judged by more than a demo:

  • Checkout time at the 99th percentile, and the slowest reply, from the server’s logs. The deadline should cap the worst case; accept nothing over 3 seconds.
  • Charges per checkoutId, from the processor’s export. It should be exactly one. Any checkoutId with two is a duplicate charge, and the number before the change is your baseline.
  • Orders held for review, and how many were later approved. This is the price of failing closed. If a person cannot clear them in a day, the limit is wrong.
  • Pending charges older than five minutes. An outcome that stays unknown is a bug, not a state.

This page did not run the checkout for real buyers, so it has no numbers to give you. How to choose the right measure is Defining success; how to take the before picture is Baseline before you change.

08 / Ask for it

Two prompts, two builds, seven broken dependencies.

We sent two agents the same request for this checkout at the same time, both running Claude Sonnet. One prompt described the product and the two APIs. The other added one paragraph: list every call to another system; for each, decide what happens when it is slow, unreachable, unexpected, repeated, or half done; implement each decision; write the list to FAILURE-MODES.md. It named no deadline, no key, and no policy. Then a script started each build against a fake fraud check and a fake processor and broke them, one question at a time.

What the checker found, run 2026-09-23
What the checker brokePlain promptArchitecture prompt
An ordinary $50 purchase200 with a card after 13 ms. 1 charge.200 with a card after 13 ms. 1 charge.
The fraud check never answersNo reply within the checker’s 40.0 s. 0 charges.503 after 10.3 s. 0 charges.
The fraud check refuses connections502 after 6 ms. 0 charges.503 after 257 ms. 0 charges.
The fraud check answers without a score502 after 13 ms. 0 charges.503 after 9 ms. 0 charges.
The same checkout sent twice200 with a card after 11 ms; then 200 with a card after 1 ms. 1 charge.200 with a card after 12 ms; then 200 with a card after 0 ms. 1 charge.
A double tap while the fraud check takes 3 s200 with a card after 3.0 s; then 200 with a card after 3.0 s. 1 charge.200 with a card after 3.0 s; then 200 with a card after 3.0 s. 1 charge.
The charge’s reply is lost, then the app retries502 after 12 ms; then 200 with a card after 3 ms. 1 charge.200 with a card after 314 ms; then 200 with a card after 0 ms. 1 charge.

Both builds handled almost everything. Both refused to charge without a valid score, both answered a second tap with the first reply, and both sent the checkoutId as the processor’s idempotency key, so no question produced a second charge.

They could, because the plain prompt already carried those decisions in its product sentences. “The buyer’s app creates the checkoutId when the payment screen opens” tells a careful reader that a retry reuses the id. “It accepts an optional Idempotency-Key header” is the processor’s documentation, and both agents used it.

The one thing no sentence said was how long to wait. The plain build asks the fraud check with a bare fetch, and when the fraud check stopped answering, the checker gave up at 40 seconds with the build still waiting.

server.ts · plain prompt
const res = await fetch(new URL(path, baseUrl), {
  method: "POST",
  headers: { "content-type": "application/json", ...headers },
  body: JSON.stringify(payload),
});
server.ts · architecture prompt
const maxAttempts = 2; // fraud check is stateless/side-effect-free, safe to retry once
const timeoutMs = 5000;

for (let attempt = 1; attempt <= maxAttempts; attempt++) {
  try {
    const res = await fetchWithTimeout(
      `${FRAUD_URL}/score`,
      {
        method: "POST",
        headers: { "content-type": "application/json" },
        body: JSON.stringify({ amountCents, email }),
      },
      timeoutMs,
    );

The architecture build did set a deadline: 5 seconds, then one retry. The buyer waited 10.3 seconds. That is the second finding: a deadline per attempt is not a deadline per checkout. Read from its source, a bad day for both vendors could hold a buyer for about 35 seconds: two 5-second fraud attempts, then three 8-second charge attempts, with pauses between.

So the prompt line the runs showed was missing is not about the five questions. The agent asked those well. It is the number the product owns: the buyer hears back within n seconds, however many attempts that allows.

How the runs were made and checkedOne run each, recorded as written
  • Both agents received the prompts word for word, in fresh contexts, in the same message. Neither was told about the other, this lesson, or the checker. The only differences were the Architecture paragraph and the output folder.
  • The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples, including the architecture build’s FAILURE-MODES.md. The checker restores them into a temporary folder and starts a fresh server for every question.
  • The fake processor honors the idempotency key, as real processors do. Stripe’s documentation says it “works by saving the resulting status code and body of the first request made for any given idempotency key” (Idempotent requests).
  • The architecture build chose to fail closed for every amount: no score, no charge. That is a sound policy the prompt left to it. This lesson’s example fails open at $25 and under, to show that the choice is a business decision with a limit attached.
  • Neither build survives a restart between the charge and the card, because both keep cards in memory as the prompt allowed. The architecture agent wrote that limit down; the checker does not test it.
  • This is one sample of each prompt, not a measurement of a model. Another run could land differently. The transcript audit and every number are in the run notes beside the examples.

09 / Hold it there

Make the missing deadline a failing check.

An inventory written once drifts the moment someone adds a fourth call. Three kinds of check keep the answers in place.

  1. The door the platform gives you

    Neither runtime closes it for you. In TypeScript, a deadline is a signal you pass: AbortSignal.timeout(ms), which aborts with a TimeoutError (MDN). In Go, it is context.WithTimeout on the request, or a nonzero http.Client.Timeout. On the other side of the charge, the door is the processor’s idempotency key; the checkout only has to send it.

  2. A rule an agent cannot argue with

    “Every outbound call has a deadline” is checkable. This ESLint rule fails any fetch without a signal. We ran it against both recorded builds and this lesson’s checkout: it flagged one line, the plain build’s bare fetch, and passed the rest. It checks that a signal is there, not that its number is sensible. Enforcement layer runs rules like this on every change.

    eslint.config.js
    // eslint.config.js (the rule only; a TypeScript project also sets its parser)
    // Run on 23 September 2026 with ESLint 10.10.0 against both recorded builds and the
    // lesson's checkout.ts. Output: deadline-rule-output.txt.
    export default [
    	{
    		files: ['server/**/*.ts', 'server.ts'],
    		rules: {
    			'no-restricted-syntax': [
    				'error',
    				{
    					selector: "CallExpression[callee.name='fetch']:not(:has(Property[key.name='signal']))",
    					message: 'Every outbound call needs a deadline: pass signal: AbortSignal.timeout(ms).'
    				}
    			]
    		}
    	}
    ];
    
  3. A check on what actually happens

    A lint rule sees code. It cannot see a processor that charges and drops the connection. So make each row of the inventory happen: a fake that loses one reply, a fake that never answers, a fake that answers the wrong shape. The checker in section 08 does exactly that, and so do this lesson’s shared cases.

    check-runs.mjs
    if (processor.mode === 'drop-first' && processor.calls === 1) {
    	request.socket.destroy(); // charged, and the reply never arrives
    	return;
    }
Build UIs?Every fetch you write has these failure modes. The Pay button is where you own them.

Where it already is in your components

fetch has three endings, and only one of them throws. A 500 resolves like a 200, so response.ok is a question you have to ask. A network failure rejects. And a reply that never comes waits, unless you pass a signal. Every save button you have written already sits on those three endings, whether or not it handles them.

Data libraries name the difference for you. TanStack Query gives every query a status (is there data, or an error?) and a separate fetchStatus (is a request in flight, or paused because the device is offline?), because “no data yet” and “still waiting” are different failure modes (Queries guide).

When you have to own it

The Pay button is the buyer’s half of this lesson’s inventory. It creates one checkoutId when the screen opens and sends the same one on every tap, so the server’s ledger can recognize a second tap. It stops waiting at 10 seconds, and then it does not say “failed”, because the card may have been charged. It says “we did not hear back” and offers to check again, which is safe for exactly the reason the server made it safe.

A 202 is not an error either. “We’re checking this order” and “payment pending” are real answers from the server, and the screen shows them as such.

Saving a note. The three ways a fetch can end, handled one by one: a reply that is ok, a reply that is not, and no reply within 5 seconds.

ReactAlready in your code
SaveNote.tsx
import { useState } from 'react';

// fetch has three ways to end, and only one of them throws:
// a reply that is ok, a reply that is not ok (a 500 still resolves), and no reply at all.
type Sent =
	{ kind: 'reply'; status: number; body: unknown } | { kind: 'no-reply'; timedOut: boolean };

export async function send(url: string, body: unknown, timeoutMs: number): Promise<Sent> {
	try {
		const response = await fetch(url, {
			method: 'POST',
			headers: { 'content-type': 'application/json' },
			body: JSON.stringify(body),
			// Without a signal, fetch waits as long as the connection stays open.
			signal: AbortSignal.timeout(timeoutMs)
		});
		return {
			kind: 'reply',
			status: response.status,
			body: await response.json().catch(() => null)
		};
	} catch (error) {
		const timedOut = error instanceof DOMException && error.name === 'TimeoutError';
		return { kind: 'no-reply', timedOut };
	}
}

export default function SaveNote({ text }: { text: string }) {
	const [message, setMessage] = useState('');

	async function save() {
		const sent = await send('/api/notes', { text }, 5000);
		if (sent.kind === 'no-reply')
			setMessage(sent.timedOut ? 'Still waiting after 5 s. Try again.' : 'Offline. Try again.');
		else if (sent.status >= 400) setMessage(`Not saved (${sent.status}).`);
		else setMessage('Saved.');
	}

	return (
		<p>
			<button onClick={save}>Save</button> <span role="status">{message}</span>
		</p>
	);
}

10 / Make the call

Ask all five questions everywhere. Pay for policies where a row can hurt.

The inventory is cheap: a table and a fake per call. The policies are what cost something: a ledger, a key, a review queue, a pending state the screen has to show. Build them where a row can lose money, data, or trust. Keep the hopeful version where a failure costs a person one more click and a repeat does no harm: a read-only admin page, a report you run by hand, a preview.

Reopen the inventory when a call starts charging, sending, or writing somewhere; when a new caller arrives with retries of its own, such as a kiosk or a mobile app; and when a vendor changes its API. Those are the moments a row changes.

Take it with you

Explain it without saying “failure mode”: “For each thing my code calls, I asked what happens if it is slow, down, wrong, sent twice, or half done. I decided each answer, and I wrote a test that makes each one happen.” Then open the last feature you built with an AI, list the calls it makes to other systems, and find the deadline on each one, or the place where there is none.

Paste into your next prompt, and fill in the blanks

Before writing code, list every call this feature makes to another system.
For each call, decide what happens when it is:
- slow: stop waiting after <deadline>; then <policy>. The whole request
  answers within <budget>, however many attempts that allows.
- unreachable: <policy>.
- wrong: accept only <valid reply>; treat anything else as <policy>.
- sent twice: <side effect> happens once per <key>; a repeat returns the stored result.
- half done: record each step before the call it guards, so a retry resumes
  from <step> instead of starting over.
Write the list to FAILURE-MODES.md, and add a test for each row that makes
that failure happen with a fake.
Connections to follow nextRelated lessons

Take the checkout into your editor. Add a fourth call, emailing the code to the buyer, and write its five rows before you write its code.

Back to architecture →