01 / The prompt
“When someone pledges, save it and publish a PledgeRecorded message.”
A donations platform. A donor pledges 5,000 cents; a sponsor’s matching fund matches it, and a receipts service emails the donor. Both of them learn about pledges from a message broker. The obvious handler saves the pledge and then publishes the message. It passes every test, because no test stops the process between the two lines.
Production does. A deploy, an out-of-memory kill, or a broker that is restarting lands between the save and the publish, and the pledge exists with nobody told: no match, no receipt. Put the publish first and the same crash does the opposite: in the lesson’s example, the sponsor matches 5000 cents for a pledge the platform never recorded. The database’s transaction cannot help, because the broker is not in it.
The brief never asked the question: if the process dies between saving and announcing, which one survives, and who finishes the other?
02 / Name the shape
Write the pledge and its message together. Let a relay send the message.
A transactional outbox is a table in the service’s own database that holds messages waiting to go out. The request writes the business row and the outbox row in one transaction, so they exist together or not at all. A relay reads pending rows, publishes each one, and marks it sent. The request never talks to the broker, so a broker outage delays announcements instead of losing them.
The change and its message commit together. The relay publishes, then marks. Every message may arrive twice, so every receiver skips an id it has already handled.
Who owns what:
| Part | Owns | Promises |
|---|---|---|
| The request | The pledge and its outbox row | Both rows in one transaction, or neither |
| The outbox table | Messages not yet sent | A row stays pending until the broker has taken it |
| The relay | Getting pending rows to the broker | In order, retried until accepted, marked only after |
| Each receiver | What it does with a message | A message id it has handled changes nothing |
Words to put in a prompt or a review
- Dual write
- Writing to two systems one after the other and hoping nothing happens in between.
- Outbox row
- The message, with its id, stored in the same database and transaction as the change it announces.
- Relay
- The loop that publishes pending rows. Also called a message relay or outbox publisher.
- Pending age
- How long the oldest unsent row has waited. It is the one number to alert on.
- At-least-once
- Every message is delivered, some of them twice. It is what an outbox can promise.
Why not publish inside the database transaction?Two systems
Calling the broker between BEGIN and COMMIT does not make it part
of the transaction. If the commit fails after the publish, the message is out and the pledge
is not; the same ghost, a few lines later. It also holds database locks while waiting on the
network. The outbox works because both writes go to one system that can commit them together.
03 / A crash between the writes
Same pledges, same crash. Which platform still tells the truth?
Each column runs one of the lesson’s platforms on the same steps, and audits the database against the broker: every recorded pledge announced, nothing announced that was never recorded. Watch four situations, then open Try it and crash them yourself.
A pledge, a crash, and the announcement
Save, then publish
Pledge reply: crashed
Database
- pledge p-1
Broker
- no messages
Matching fund: 0 cents
Save the message with the pledge
Pledge reply: crashed
Database
- pledge p-1
- outbox msg-1 · pending
Broker
- no messages
Matching fund: 0 cents
Ada pledges 5000 cents, and the process dies after its first write.
Ada pledges 5,000 cents and the process dies straight after its first write. Saved alone, the pledge exists and nobody hears of it: no match, no receipt. Saved with its message, the relay finds the row and announces it.
Reduced motion: choose a scene to see its completed state.
Read this scene
Ada pledges 5,000 cents and the process dies straight after its first write. Saved alone, the pledge exists and nobody hears of it: no match, no receipt. Saved with its message, the relay finds the row and announces it.
Save, then publish: pledges p-1, broker messages none, matched 0 cents.
Save the message with the pledge: pledges p-1, broker messages none, matched 0 cents.
Watch restarts the story when you come back. Step through shows where each chapter ends. Try it starts a new platform whenever you change the design or press Reset.
04 / Read the shape
One transaction for two rows, and a relay that marks only what the broker took.
Basic form records a pledge. In the wild is the relay. At the call site shows who calls each. Notice that no line in the request mentions the broker.
Recording a pledge: the pledge row and its outbox row are written in one transaction, with the message id fixed there. The request never calls the broker, so a broker outage cannot fail it.
/**
* The pledge and its message are one transaction: both rows, or neither.
* The message id is fixed here, so every later attempt to publish it
* carries the same id. The request never talks to the broker.
*/
protected recordPledge(request: PledgeRequest): Reply {
const pledge = this.pledgeOf(request);
this.transaction((tables) => {
tables.pledges.push(pledge);
const id = `msg-${tables.outbox.length + 1}`;
tables.outbox.push({ id, pledge: pledge.id, amount: pledge.amount, published: false });
});
this.crashAt(request, 'after-commit');
return { status: 'recorded' };
} // Record writes the pledge and its message in one transaction: both rows, or
// neither. The message id is fixed here, so every later attempt to publish it
// carries the same id. The request never talks to the broker.
func (p *OutboxPlatform) Record(r PledgeRequest) string {
pl := pledgeOf(r)
p.transaction(func(t *Tables) error {
t.Pledges = append(t.Pledges, pl)
id := "msg-" + itoa(len(t.Outbox)+1)
t.Outbox = append(t.Outbox, OutboxRow{Message{id, pl.ID, pl.Amount}, false})
return nil
})
return outcome("recorded", crashAt(r, "after-commit"))
} The two dual writesSave, then publish; publish, then save
Both look correct, and both are, until the process dies between their two lines. The first loses the announcement; the second announces a pledge that does not exist.
/** Save, then publish: a crash in between records a pledge nobody hears about. */
export class SaveThenPublish extends Platform {
protected recordPledge(request: PledgeRequest): Reply {
const pledge = this.pledgeOf(request);
this.transaction((tables) => tables.pledges.push(pledge));
this.crashAt(request, 'after-save');
const id = `msg-${this.broker.messages.length + 1}`;
const ok = this.broker.publish({ id, pledge: pledge.id, amount: pledge.amount });
return { status: ok ? 'recorded' : 'failed' };
}
}
/** Publish, then save: a crash in between announces a pledge that was never recorded. */
export class PublishThenSave extends Platform {
protected recordPledge(request: PledgeRequest): Reply {
const pledge = this.pledgeOf(request);
const id = `msg-${this.broker.messages.length + 1}`;
if (!this.broker.publish({ id, pledge: pledge.id, amount: pledge.amount }))
return { status: 'failed' };
this.crashAt(request, 'after-publish');
this.transaction((tables) => tables.pledges.push(pledge));
return { status: 'recorded' };
}
} The behavior these examples promiseChecked by 18 shared scenarios
- Save, then publish: a crash after the save, or a broker that is down, leaves a pledge that was never announced, and nothing tries again.
- Publish, then save: a crash after the publish leaves an announcement, and a match, for a pledge that was never recorded.
- The outbox: every recorded pledge is announced once the relay runs, whatever crashed and whenever the broker was down. A relay that stops after publishing sends the message again with the same id, and the matching fund matches it once.
Every expectation was generated by a separate model written from the contract in the examples’ README, not copied from either implementation, and it is kept beside the examples.
Reading the TypeScriptA transaction as a copy
transaction runs the work on a structuredClone of the tables
and swaps it in only if the work finishes, so a crash inside it leaves nothing behind,
the way a rolled-back transaction does. A crash is a thrown Crash, caught
by record and reported as crashed.
Reading the GoThe same copy, and an error
Go copies the two slices, runs the work, and assigns the copy back only when the work
returns no error. A crash is errCrash, returned by the step where the
process dies. The three designs share a base struct through embedding.
Run it yourselfNo dependencies
Save the complete files at the paths in their banners. Then run node --experimental-strip-types run.ts (Node 22.18 or later), or go run . in the Go folder. Both print:
save-then-publish · crash between the writes: pledges [p-1], messages [], matched 0 -> never-announced:p-1 publish-then-save · crash between the writes: pledges [], messages [msg-1], matched 5000 -> ghost:p-1 outbox · crash after the commit: pledges [p-1], messages [msg-1], matched 5000 -> consistent save-then-publish · broker down, then up: pledges [p-1], messages [], matched 0 -> never-announced:p-1 outbox · broker down, then up: pledges [p-1], messages [msg-1], matched 5000 -> consistent outbox · relay stops, runs again: pledges [p-1], messages [msg-1 msg-1], matched 5000 -> consistent
05 / Review the agent’s diff
“Marking first fixes it.”
A donor got two receipts, and the agent found why. Read what its fix does to a crash.
06 / How it fails
Every failure here is a crash in the gap: the question is what was written down.
Each row except the last two is a shared scenario the tests run.
| What happens | Save, then publish | Publish, then save | The outbox |
|---|---|---|---|
| The process dies after its first write | A pledge nobody hears about. | A match for a pledge that does not exist. | Both rows committed; the relay announces it. |
| The broker is down for a minute | The pledge is saved, the publish fails, and nothing retries. | The request fails; the donor tries again. | The request succeeds; the row waits for the broker. |
| The relay stops after a publish, before the mark | No relay. | The message goes twice with one id; the fund matches once. | |
| A receiver does not check ids | Rarely matters: each message is sent at most once. | A second receipt, or a second match. Not modeled in the example. | |
| The relay is stuck | Not applicable. | Pending rows pile up and announcements stop, silently unless someone watches pending age. Not modeled in the example. | |
The receiving side has its own lessons: Idempotency and at-least-once explains why a message must be safe to receive twice, and Receiving webhooks keeps a receipt table for exactly that.
07 / Is it worth it?
A table and a relay, against two lines that are almost always fine.
| Change | Dual write | The outbox |
|---|---|---|
| A second entry point: pledges imported from a partner’s CSV | The import must remember to publish, and has its own gap. | The import writes the same two rows; the relay does the rest. |
| A new receiver: a leaderboard | No change. | No change. |
| A new rule: a canceled pledge is announced too | A second dual write with the same gap. | A second message type in the same transaction as the cancel. |
| The broker is replaced | Every call site changes. | The relay changes; requests do not. |
The costs are real: a table that grows until something deletes published rows, a relay to run and watch, announcements that lag by one relay turn, and receivers that must handle a copy. When a lost or phantom message costs nothing, say a cache hint or an analytics ping, the dual write is fine. When it is money or a promise to a person, it is not.
Measure before and after:
- Records with no message, and messages with no record, from a daily reconciliation. With a dual write, this is the number that is not zero.
- Age of the oldest pending row, which is how late announcements are running, and duplicate messages per day seen by receivers.
- Outbox table size, against your cleanup policy.
This lesson did not measure a real platform, and gives no numbers.
08 / Ask for it
One brief, two prompts.
Two agents running Claude Sonnet each got the brief from section 01. One prompt added an Outbox block: the pledge and its message in one transaction, a relay that publishes in order and marks after the broker accepts, a message id fixed when the row is written, and a broker that never slows a pledge. A script ran both builds against its own broker, which it took down, slowed, and made hold a message without answering, and it killed the servers along the way.
| Question | Plain prompt | Outbox prompt |
|---|---|---|
| Ten pledges, one after another | 10 of 10 announced, out of order; replies in 4 ms or less | 10 of 10 announced, in order; replies in 2 ms or less |
| The broker down for 20 seconds | Every pledge answered 201; 5 of 5 announced 1010 ms after it came back, out of order | Every pledge answered 201; 5 of 5 announced 202 ms after it came back, in order |
| A broker that takes 8 seconds to answer | Replies in 2 ms or less; 3 of 3 announced, 7 messages sent for 3 pledges | Replies in 2 ms or less; 3 of 3 announced, 3 messages sent for 3 pledges |
| Server killed while the broker holds a publish | 1 of 1 announced; sent again after the restart with the same id (1 copy) | 1 of 1 announced; sent again after the restart with the same id (1 copy) |
| Server killed with five pledges pending | 5 of 5 announced | 5 of 5 announced |
| Twenty pledges at once | 20 of 20 announced; one id per pledge | 20 of 20 announced; one id per pledge |
| Its own tests | 14 of 14 pass | 19 of 19 pass |
Both agents built an outbox. The brief said “a pledge must never be announced unless it was recorded, and a recorded pledge must be announced”, and that sentence was enough: the plain agent wrote the pledge and its message in one transaction, marked a row only after a 202, and resumed after a restart. No pledge was lost and none was announced without being recorded, in either build, under any question.
The difference was order, and the brief never mentioned it. The plain build gives each outbox row a random UUID and publishes pending rows sorted by that id, so pledges reach the broker in an order nobody chose. For a matching fund that adds up totals, that is harmless. For a receiver that sees “pledge canceled” before “pledge recorded”, it is not.
const messageId = randomUUID();
…
.prepare('SELECT id, pledge_id, payload FROM outbox WHERE published = 0 ORDER BY id LIMIT ?') The outbox build keeps a separate sequence number for the order and the message id for deduplication, and publishes one row at a time, oldest first.
// seq is the durable publish order: rows are relayed in ascending seq
// order. id is the message id that goes out to the broker; it is fixed
// the moment the row is written, so republishing the same row (after a
// crash, or a lost 202) always carries the same id and downstream
// consumers can dedupe on it.
db.exec(`
CREATE TABLE IF NOT EXISTS outbox (
seq INTEGER PRIMARY KEY AUTOINCREMENT,
id TEXT NOT NULL UNIQUE,
topic TEXT NOT NULL, The missing line is one sentence: publish pending rows in the order they were written, by a sequence, not by the message id. The rest of the outbox block restated what the plain agent already did. A brief that states the guarantee gets most of the mechanism; it does not get the properties nobody wrote down.
The slow broker showed the other unstated number. The plain build gives up on a publish after 5 seconds and tries again, so an 8-second broker got several copies of each message; the outbox build waits 10 seconds and sent each once. The copies carried the same id, so nothing downstream was wrong, but the timeout against the broker’s real latency is a number to ask for.
How the runs were made and checkedTwo builds, recorded as written
- Both agents were launched at the same time from empty folders; neither was told about the other, the lesson, or the checker.
- Both builds are kept byte for byte with checksums. For every question the checker restores a build into a fresh folder with its own database, starts the server, and runs its own broker.
- The checker ran twice; the first covered the plain build alone while the outbox agent was still working, and gave the same answers apart from timings.
- Both agents wrote a few scratch files in
/tmpduring their own checks, against the prompt, and deleted them. Neither stopped a process by name or pattern. - One run of each prompt is a sample, not a measurement of the model.
09 / Hold it there
An outbox breaks when someone adds a publish outside it. Three checks notice.
Only the relay imports the broker client
The moment a request handler calls the broker directly, the dual write is back. An import rule that allows the broker client in the relay and nowhere else keeps it out, the way Enforcement layer enforces import rules.
Tests that stop the process between the steps
A test that records a pledge and waits proves nothing about the gap. The shared scenarios crash after each write, take the broker down, and stop the relay between its publish and its mark; a check that the pledge and its outbox row are written by one transaction catches the rest.
Alert on pending age, and reconcile
Alert when the oldest pending row is older than a relay turn or two. Once a day, compare pledges with the messages the matching fund received; with an outbox the difference should be the rows still pending, and nothing else.
There is no frontend version of this lesson: the outbox sits between a service’s database and its broker, which a browser never sees. The browser’s own version, a change and its unsent request kept together in local storage until the server takes them, is in Local-first and sync.
10 / Make the call
If the message matters, it goes in the transaction.
Use an outbox whenever a change must be followed by a message that another system acts on: money, emails to people, stock, access. Keep the dual write for messages whose loss nobody would notice. Reopen the design when the database already publishes its own changes; then Change data capture can play the relay.
Take it with you
Explain it without saying “outbox”: “When we take a pledge, we write the pledge and a note saying ‘tell the sponsor about this’ on the same page, in one go. Someone reads the notes and passes each one on, ticking it off after. If they are interrupted, they pass it on again, and the sponsor ignores a note it has already seen.” Then find a place in your own code that saves something and then sends a message, and ask what happens if the process dies on the line between them.
Paste into your next prompt, and fill in the blanks
When [a pledge] is recorded, [the matching fund and receipts service] must hear about it through [the broker]. Write the [pledge] and its message to an outbox table in the same database transaction; the request never calls the broker. Give each message its id when the row is written, and publish it with that id every time. A relay in the same service publishes pending rows in the order they were written, by a sequence number, marks each one only after the broker accepts it, and keeps trying while the broker is down; a restart resumes from the rows still pending. Receivers may get a message twice: they skip an id they have already handled. Alert when the oldest pending row is older than [one minute], and delete published rows after [seven days].
Connections to follow nextRelated lessons
- Event-driven architecture is what the messages are for.
- Consistency without a shared transaction is the same gap between modules, finished forward or undone.
- Sagas and compensation chains several of these steps, each announced through an outbox.
- Change data capture reads the database’s own log instead of a table the code writes.
- Idempotency and at-least-once is the receiving side’s half of the bargain.