01 / The prompt
“Build the checkout for our gift cards.”
You ask an agent for the online checkout of a café chain’s gift cards. Ask a third-party fraud check whether to go ahead, charge the card through the card processor, create the gift card, show its code. What comes back works. You buy a $50 card, the fraud check answers in a fifth of a second, and the code appears. Every test passes.
Nothing in those first ten minutes tells you what the checkout does on the day the fraud
check takes 30 seconds to answer. The default is not reassuring. Go’s HTTP client says it in
one line: “A Timeout of zero means no timeout” (net/http, Client), and zero is what you get unless you set it. Node’s built-in fetch comes
from undici, which waits up to 300 seconds for a reply’s headers before it gives up (headersTimeout, default 300e3). Both defaults were checked on 23 September 2026.
So the checkout waits as long as the fraud check does. Meanwhile the buyer is looking at a spinner, and the button next to it still says Pay.
The prompt never said how long to wait, what to do without a score, or what a second tap means. The agent had to pick something for each, and a demo cannot show you what it picked.
02 / Name the move
Five questions for every call you do not control.
A failure mode is one specific way a call can go wrong, together with what your system does when it happens. Thinking in failure modes is the habit of listing every call that crosses a boundary and asking the same five questions of each one before the code ships.
What if it is slow? Down? Wrong? Sent twice? Half done? Give every answer a policy, and a test that makes it happen.
The gift-card checkout makes three calls. Here is who runs the other side of each, and what the checkout cannot know about it.
| Call | Who runs the other side | What the checkout cannot know |
|---|---|---|
| Score the purchase | The fraud vendor | Whether it will answer, how soon, and in which shape. |
| Charge the card | The card processor | Whether a charge happened, when no reply comes back. |
| Write the card | The café, in its own database | Nothing, until the process stops between the charge and this write. |
Words to put in a prompt or a review
- Deadline
- How long you will wait for a call before you decide without it.
- Fail open, fail closed
- Without an answer, go ahead anyway, or stop. Name which, and when.
- Idempotency key
- A name for one operation, so the other side can recognize a repeat of it.
- Outcome unknown
- No reply after a side effect. It may have happened. Say so; do not guess.
- Ledger
- Your record of each step, written before the call it guards.
- Failure inventory
- The table of calls, conditions, policies, and tests. The thing you hand over.
Is this FMEA?The engineering version, and why this one is smaller
Failure mode and effects analysis comes from reliability engineering. It lists how each component can fail and what each failure does to the whole system, and it usually ranks them. The habit here is the same idea sized for one request path: the rows are the calls your code makes, the five questions are fixed, and the output is a table of policies with a test behind each row. Rank them if you have many; for three calls, write them all.
03 / Follow one checkout
Watch the same bad day hit two checkouts.
First the hopeful build, on the day the fraud check takes 30 seconds. Then the inventoried build on the same day. Then the inventoried build again, when the processor charges the card and the reply never arrives. Step through at your own pace, or open Try it and break the calls yourself.
What does a checkout do while a dependency is slow?
Hopeful build
Buyer’s app
Paying $50…
Checkout
- This tap has waited
- 0 ms
Charges 0Cards 0
-
Fraud check
Not asked yet
-
Card processor
Not asked yet
-
Gift-card store
Not asked yet
Buyer → checkoutPOST /checkout {"checkoutId":"gc-1","amountCents":5000}
The buyer taps Pay.
The fraud check is having a bad day. The hopeful build asks it with no deadline, because nothing in the prompt mentioned one.
Reduced motion: choose a scene to see its completed state.
Read this scene
The fraud check is having a bad day. The hopeful build asks it with no deadline, because nothing in the prompt mentioned one.
Hopeful build. The buyer’s screen: Paying $50…. This tap has waited 0 ms. Charges at the processor: 0. Cards issued: 0. Fraud check: Not asked yet. Card processor: Not asked yet. Gift-card store: Not asked yet.
On the wire: POST /checkout {"checkoutId":"gc-1","amountCents":5000}
Watch restarts the story when you come back. Step through keeps your step. Try it builds a fresh world for both checkouts every time you run it.
04 / Read it in code
A deadline, a shape check, a key, and a ledger.
Basic form answers three questions for one call: slow, down, wrong. In the wild answers the other two for the whole checkout: twice and half done. At the call site the deadline stops being a number and becomes something the network layer enforces.
Notice what the checkout does not do: it never guesses. No score is its own case with its own policy. No reply from the processor is “pending”, not “failed”, because a failure you report wrongly makes the buyer pay again.
The fraud call with the first three questions answered: a deadline for slow, a status check for down, and a shape check for wrong. Anything that is not a whole-number score from 0 to 100 is not a score.
export const FRAUD_DEADLINE_MS = 2000;
export type Verdict =
{ kind: 'scored'; score: number } | { kind: 'timeout' | 'unavailable' | 'invalid' };
/** Ask the fraud check, but only for so long, and believe only a reply that makes sense. */
export async function scoreWithin(deps: Deps, amountCents: number): Promise<Verdict> {
const reply = await deps.score(amountCents, FRAUD_DEADLINE_MS);
if (reply.timedOut) return { kind: 'timeout' };
if (reply.status !== 200) return { kind: 'unavailable' };
const score = field(reply.body, 'score');
if (typeof score !== 'number' || !Number.isInteger(score) || score < 0 || score > 100)
return { kind: 'invalid' };
return { kind: 'scored', score };
} const FraudDeadline = 2 * time.Second
// Verdict is "scored", "timeout", "unavailable", or "invalid".
type Verdict struct {
Kind string
Score int
}
// ScoreWithin asks the fraud check, but only for so long, and believes only a reply that makes sense.
func ScoreWithin(deps Deps, amountCents int) Verdict {
reply := deps.Score(amountCents, FraudDeadline)
if reply.TimedOut {
return Verdict{Kind: "timeout"}
}
if reply.Status != http.StatusOK {
return Verdict{Kind: "unavailable"}
}
var body struct {
Score *float64 `json:"score"`
}
if json.Unmarshal([]byte(reply.Body), &body) != nil || body.Score == nil {
return Verdict{Kind: "invalid"}
}
score := *body.Score
if score != float64(int(score)) || score < 0 || score > 100 {
return Verdict{Kind: "invalid"}
}
return Verdict{Kind: "scored", Score: int(score)}
} The behavior these examples promiseChecked by 80 shared cases in TypeScript and Go
- The fraud call has a 2-second deadline. A reply counts as a score only if it is a 200
whose
scoreis a whole number from 0 to 100. A score of 70 or more declines. - Without a usable score, a card of $25 or less is charged; a larger one is held for review with a 202 and nothing is charged.
- The charge is sent with the checkoutId as its idempotency key and a 5-second deadline. Any reply without a charge id is 202 pending, and the ledger stays at charging.
- A checkoutId with a final reply gets that reply again, marked as stored, with no calls. One left at charging asks the processor again. One left at charged writes the card.
- The hopeful build waits without a deadline, reads a missing score as 0, charges with no key, and keeps no ledger.
Every expectation in the shared cases was produced by a separate model written from these
rules, kept beside the examples in examples/model/, not copied from either
implementation. It covers every combination of five fraud conditions, three processor
conditions, a stop before the card, two amounts, and one or two taps.
Reading the TypeScriptunknown, a union, and AbortSignal.timeout
field returns unknown, so the score has to be proved a number
before anything compares it. Verdict is a union on kind, so
“no score” cannot be mistaken for a score of zero, which is exactly the mistake the
hopeful build’s { score = 0 } makes. AbortSignal.timeout rejects the fetch with a TimeoutError DOMException (MDN), which is how the adapter tells a timeout from a refused connection.
Reading the GoA pointer to tell missing from zero, and a context
Decoding into *float64 leaves the pointer nil when the field is missing, so
a missing score and a score of 0 stay different. The hopeful build decodes into a plain int and cannot tell them apart. context.WithTimeout puts the
deadline on the request, and errors.Is(err, context.DeadlineExceeded) says
it was the deadline that ended it. A deadline of 0 means none, the same convention as http.Client.
Run it yourselfNo dependencies
Copy the complete TypeScript file and run node --experimental-strip-types checkout.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:
module heyrian.dev/lessons/failure-modes
go 1.23
gc-1 $50, fraud answers: 200 issued GC-0001 after 500 ms | charges 1, cards 1 gc-2 $50, fraud slow: 202 review after 2000 ms, then 202 review (stored) after 0 ms | charges 0, cards 0 gc-3 $20, fraud slow: 200 issued GC-0001 after 2320 ms | charges 1, cards 1 gc-4 $50, charge reply lost: 202 pending after 480 ms, then 200 issued GC-0001 after 320 ms | charges 1, cards 1 gc-5 $50, stopped before the card: 502 stopped after 480 ms, then 200 issued GC-0001 after 20 ms | charges 1, cards 1
The times are the simulated world’s clock, not a measurement: each fake says how long it would take, and the clock adds it up.
05 / Review the agent’s diff
“Checkout no longer hangs when a dependency is slow.”
The complaint was real, and adding timeouts is the right instinct. Read what else changed before you decide. The question to ask of each call is not “can this fail?” but “what happens if it is repeated?”
06 / How it fails
The inventory: every call, every question, every row tested.
This is the artifact the procedure produces. Each row is a call and a condition, what the buyer sees, and what the inventoried checkout does. The last column comes from the shared cases, and the two-tabs row from a native test in each language, so each row is a test that fails if the policy changes.
| Call and condition | What the buyer sees | What the checkout does | Hopeful → inventoried |
|---|---|---|---|
| Fraud checkSlow: takes 30 s | “We’re checking this order. No charge yet.” after 2 s. A $20 card goes through. | Stops waiting at 2 s. With no usable score, $25 and under goes through and anything more waits for a person. | 2 charges, 2 cards → 0 charges, 0 cards |
| Fraud checkDown: refuses the connection | The same as slow, sooner. | The same policy. It cannot tell a dead vendor from a slow one, and does not need to. | 0 charges, 0 cards → 0 charges, 0 cards |
| Fraud checkWrong: answers {"risk":"high"} | “We’re checking this order.” | Believes only a whole-number score from 0 to 100. Anything else counts as no score. | 1 charge, 1 card → 0 charges, 0 cards |
| Fraud checkAnswers a high score, 91 | “This payment was declined.” | Declines and stores the answer. Both builds agree here. | 0 charges, 0 cards → 0 charges, 0 cards |
| ChargeDown: refuses the connection | “Payment pending. Don’t pay again.” | Replies 202 pending and keeps the ledger at charging, so the next tap asks the processor again under the same key. | 0 charges, 0 cards → 0 charges, 0 cards |
| ChargeHalf done: charged, reply lost | “Payment pending,” then the card on the next tap. | Retries with the same key. The processor returns the charge it already made. | 2 charges, 1 card → 1 charge, 1 card |
| Whole checkoutTwice: the buyer taps Pay again | The same reply both times. | Answers from the ledger without calling anything. | 2 charges, 2 cards → 1 charge, 1 card |
| Whole checkoutConflicting: two open tabs pay at once | The same card code in both tabs. | Both requests carry the same checkout id, so the charge and the card are made under one key: the processor and the card store return what they already made. Each tab still asks the fraud check once. | 2 charges, 2 cards → 1 charge, 1 card (native test) |
| Write the cardHalf done: the process stops after the charge | An error, then the card on the next tap. | The ledger says charged, so the retry skips the charge and writes the card. | 2 charges, 1 card → 1 charge, 1 card |
Five rows share one idea: a side effect may be repeated only if the other side can recognize the repeat. The key does that at the processor, and the ledger does it in the checkout. The rest share another: no answer is its own answer. Decide in advance what it means, instead of letting a default decide.
When a row changes nothing, as the high-score row does here, keep it. It records that you asked. When a row’s test passes on the hopeful build too, check the fake before you celebrate: a fake that never loses a reply cannot tell you what a lost reply does.
Each of these has a lesson of its own: Timeouts, deadlines, and races, Idempotency and at-least-once, and Retry, backoff, and idempotency.
07 / Is it worth it?
You pay for a ledger and a review queue. Here is what they buy.
The hopeful build is shorter and just as fast on a good day. Hold both up against the changes every checkout eventually gets.
| Change | Hopeful build | Inventoried build |
|---|---|---|
| A second client: a kiosk in the café | The kiosk’s own retries charge twice, the same way a second tap does. | The kiosk sends a checkoutId and gets the same guarantees. |
| Replace the fraud vendor | A differently shaped reply reads as a score of 0, and every card goes through. | A new shape is “no score”: held for review, and visible in the review count. |
| Change a rule: review above $10, not $25 | There is no review rule to change. | One constant. No other difference. |
| A payments team takes over the charge | Nothing written down about what a retry may do. | They inherit the inventory: each row a policy and a test. |
Before you change a checkout like this, write down what you will measure and what result you would accept, so the change is judged by more than a demo:
- Checkout time at the 99th percentile, and the slowest reply, from the server’s logs. The deadline should cap the worst case; accept nothing over 3 seconds.
- Charges per checkoutId, from the processor’s export. It should be exactly one. Any checkoutId with two is a duplicate charge, and the number before the change is your baseline.
- Orders held for review, and how many were later approved. This is the price of failing closed. If a person cannot clear them in a day, the limit is wrong.
- Pending charges older than five minutes. An outcome that stays unknown is a bug, not a state.
This page did not run the checkout for real buyers, so it has no numbers to give you. How to choose the right measure is Defining success; how to take the before picture is Baseline before you change.
08 / Ask for it
Two prompts, two builds, seven broken dependencies.
We sent two agents the same request for this checkout at the same time, both running Claude Sonnet. One prompt described the product and the two APIs. The other added one paragraph: list every call to another system; for each, decide what happens when it is slow, unreachable, unexpected, repeated, or half done; implement each decision; write the list to FAILURE-MODES.md. It named no deadline, no key, and no policy. Then a script started each build against a fake fraud check and a fake processor and broke them, one question at a time.
| What the checker broke | Plain prompt | Architecture prompt |
|---|---|---|
| An ordinary $50 purchase | 200 with a card after 13 ms. 1 charge. | 200 with a card after 13 ms. 1 charge. |
| The fraud check never answers | No reply within the checker’s 40.0 s. 0 charges. | 503 after 10.3 s. 0 charges. |
| The fraud check refuses connections | 502 after 6 ms. 0 charges. | 503 after 257 ms. 0 charges. |
| The fraud check answers without a score | 502 after 13 ms. 0 charges. | 503 after 9 ms. 0 charges. |
| The same checkout sent twice | 200 with a card after 11 ms; then 200 with a card after 1 ms. 1 charge. | 200 with a card after 12 ms; then 200 with a card after 0 ms. 1 charge. |
| A double tap while the fraud check takes 3 s | 200 with a card after 3.0 s; then 200 with a card after 3.0 s. 1 charge. | 200 with a card after 3.0 s; then 200 with a card after 3.0 s. 1 charge. |
| The charge’s reply is lost, then the app retries | 502 after 12 ms; then 200 with a card after 3 ms. 1 charge. | 200 with a card after 314 ms; then 200 with a card after 0 ms. 1 charge. |
Both builds handled almost everything. Both refused to charge without a valid score, both answered a second tap with the first reply, and both sent the checkoutId as the processor’s idempotency key, so no question produced a second charge.
They could, because the plain prompt already carried those decisions in its product sentences. “The buyer’s app creates the checkoutId when the payment screen opens” tells a careful reader that a retry reuses the id. “It accepts an optional Idempotency-Key header” is the processor’s documentation, and both agents used it.
The one thing no sentence said was how long to wait. The plain build asks the fraud check
with a bare fetch, and when the fraud check stopped answering, the checker gave
up at 40 seconds with the build still waiting.
const res = await fetch(new URL(path, baseUrl), {
method: "POST",
headers: { "content-type": "application/json", ...headers },
body: JSON.stringify(payload),
}); const maxAttempts = 2; // fraud check is stateless/side-effect-free, safe to retry once
const timeoutMs = 5000;
for (let attempt = 1; attempt <= maxAttempts; attempt++) {
try {
const res = await fetchWithTimeout(
`${FRAUD_URL}/score`,
{
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({ amountCents, email }),
},
timeoutMs,
); The architecture build did set a deadline: 5 seconds, then one retry. The buyer waited 10.3 seconds. That is the second finding: a deadline per attempt is not a deadline per checkout. Read from its source, a bad day for both vendors could hold a buyer for about 35 seconds: two 5-second fraud attempts, then three 8-second charge attempts, with pauses between.
So the prompt line the runs showed was missing is not about the five questions. The agent asked those well. It is the number the product owns: the buyer hears back within n seconds, however many attempts that allows.
How the runs were made and checkedOne run each, recorded as written
- Both agents received the prompts word for word, in fresh contexts, in the same message. Neither was told about the other, this lesson, or the checker. The only differences were the Architecture paragraph and the output folder.
- The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s
examples, including the architecture build’s
FAILURE-MODES.md. The checker restores them into a temporary folder and starts a fresh server for every question. - The fake processor honors the idempotency key, as real processors do. Stripe’s documentation says it “works by saving the resulting status code and body of the first request made for any given idempotency key” (Idempotent requests).
- The architecture build chose to fail closed for every amount: no score, no charge. That is a sound policy the prompt left to it. This lesson’s example fails open at $25 and under, to show that the choice is a business decision with a limit attached.
- Neither build survives a restart between the charge and the card, because both keep cards in memory as the prompt allowed. The architecture agent wrote that limit down; the checker does not test it.
- This is one sample of each prompt, not a measurement of a model. Another run could land differently. The transcript audit and every number are in the run notes beside the examples.
09 / Hold it there
Make the missing deadline a failing check.
An inventory written once drifts the moment someone adds a fourth call. Three kinds of check keep the answers in place.
The door the platform gives you
Neither runtime closes it for you. In TypeScript, a deadline is a signal you pass:
AbortSignal.timeout(ms), which aborts with aTimeoutError(MDN). In Go, it iscontext.WithTimeouton the request, or a nonzerohttp.Client.Timeout. On the other side of the charge, the door is the processor’s idempotency key; the checkout only has to send it.A rule an agent cannot argue with
“Every outbound call has a deadline” is checkable. This ESLint rule fails any
fetchwithout asignal. We ran it against both recorded builds and this lesson’s checkout: it flagged one line, the plain build’s barefetch, and passed the rest. It checks that a signal is there, not that its number is sensible. Enforcement layer runs rules like this on every change.eslint.config.js // eslint.config.js (the rule only; a TypeScript project also sets its parser) // Run on 23 September 2026 with ESLint 10.10.0 against both recorded builds and the // lesson's checkout.ts. Output: deadline-rule-output.txt. export default [ { files: ['server/**/*.ts', 'server.ts'], rules: { 'no-restricted-syntax': [ 'error', { selector: "CallExpression[callee.name='fetch']:not(:has(Property[key.name='signal']))", message: 'Every outbound call needs a deadline: pass signal: AbortSignal.timeout(ms).' } ] } } ];A check on what actually happens
A lint rule sees code. It cannot see a processor that charges and drops the connection. So make each row of the inventory happen: a fake that loses one reply, a fake that never answers, a fake that answers the wrong shape. The checker in section 08 does exactly that, and so do this lesson’s shared cases.
check-runs.mjs if (processor.mode === 'drop-first' && processor.calls === 1) { request.socket.destroy(); // charged, and the reply never arrives return; }
Build UIs?Every fetch you write has these failure modes. The Pay button is where you own them.
Where it already is in your components
fetch has three endings, and only one of them throws. A 500 resolves like a
200, so response.ok is a question you have to ask. A network failure rejects. And
a reply that never comes waits, unless you pass a signal. Every save button you have written
already sits on those three endings, whether or not it handles them.
Data libraries name the difference for you. TanStack Query gives every query a status (is there data, or an error?) and a separate fetchStatus (is a request in flight, or paused because the device is offline?), because “no data yet” and
“still waiting” are different failure modes (Queries guide).
When you have to own it
The Pay button is the buyer’s half of this lesson’s inventory. It creates one checkoutId when the screen opens and sends the same one on every tap, so the server’s ledger can recognize a second tap. It stops waiting at 10 seconds, and then it does not say “failed”, because the card may have been charged. It says “we did not hear back” and offers to check again, which is safe for exactly the reason the server made it safe.
A 202 is not an error either. “We’re checking this order” and “payment pending” are real answers from the server, and the screen shows them as such.
Saving a note. The three ways a fetch can end, handled one by one: a reply that is ok, a reply that is not, and no reply within 5 seconds.
import { useState } from 'react';
// fetch has three ways to end, and only one of them throws:
// a reply that is ok, a reply that is not ok (a 500 still resolves), and no reply at all.
type Sent =
{ kind: 'reply'; status: number; body: unknown } | { kind: 'no-reply'; timedOut: boolean };
export async function send(url: string, body: unknown, timeoutMs: number): Promise<Sent> {
try {
const response = await fetch(url, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify(body),
// Without a signal, fetch waits as long as the connection stays open.
signal: AbortSignal.timeout(timeoutMs)
});
return {
kind: 'reply',
status: response.status,
body: await response.json().catch(() => null)
};
} catch (error) {
const timedOut = error instanceof DOMException && error.name === 'TimeoutError';
return { kind: 'no-reply', timedOut };
}
}
export default function SaveNote({ text }: { text: string }) {
const [message, setMessage] = useState('');
async function save() {
const sent = await send('/api/notes', { text }, 5000);
if (sent.kind === 'no-reply')
setMessage(sent.timedOut ? 'Still waiting after 5 s. Try again.' : 'Offline. Try again.');
else if (sent.status >= 400) setMessage(`Not saved (${sent.status}).`);
else setMessage('Saved.');
}
return (
<p>
<button onClick={save}>Save</button> <span role="status">{message}</span>
</p>
);
}
10 / Make the call
Ask all five questions everywhere. Pay for policies where a row can hurt.
The inventory is cheap: a table and a fake per call. The policies are what cost something: a ledger, a key, a review queue, a pending state the screen has to show. Build them where a row can lose money, data, or trust. Keep the hopeful version where a failure costs a person one more click and a repeat does no harm: a read-only admin page, a report you run by hand, a preview.
Reopen the inventory when a call starts charging, sending, or writing somewhere; when a new caller arrives with retries of its own, such as a kiosk or a mobile app; and when a vendor changes its API. Those are the moments a row changes.
Take it with you
Explain it without saying “failure mode”: “For each thing my code calls, I asked what happens if it is slow, down, wrong, sent twice, or half done. I decided each answer, and I wrote a test that makes each one happen.” Then open the last feature you built with an AI, list the calls it makes to other systems, and find the deadline on each one, or the place where there is none.
Paste into your next prompt, and fill in the blanks
Before writing code, list every call this feature makes to another system. For each call, decide what happens when it is: - slow: stop waiting after <deadline>; then <policy>. The whole request answers within <budget>, however many attempts that allows. - unreachable: <policy>. - wrong: accept only <valid reply>; treat anything else as <policy>. - sent twice: <side effect> happens once per <key>; a repeat returns the stored result. - half done: record each step before the call it guards, so a retry resumes from <step> instead of starting over. Write the list to FAILURE-MODES.md, and add a test for each row that makes that failure happen with a fake.
Connections to follow nextRelated lessons
- Containing failure takes one row, a slow dependency, and keeps it from slowing everything else.
- Defining success decides what the numbers in section 07 should be before you change anything.
- Timeouts, deadlines, and races goes deeper on the difference the runs found: a timeout per attempt versus a deadline for the whole request.
- Idempotency and at-least-once explains why a repeated request is the normal case, not the edge.
- Spec before code is the same move as section 08: write down what must be true, then check the build against it.