← Architecture
Events, queues, and workflows Make the handoff durable

Transactional outbox

Save the fact before you try to send it.

Saving a pledge and announcing it are two writes to two systems, and no transaction covers both. The outbox moves the announcement into the database, next to the pledge, and lets a relay carry it out afterward.

TypeScriptGoOne pledge, three designs, a relay, two recorded builds.

01 / The prompt

“When someone pledges, save it and publish a PledgeRecorded message.”

A donations platform. A donor pledges 5,000 cents; a sponsor’s matching fund matches it, and a receipts service emails the donor. Both of them learn about pledges from a message broker. The obvious handler saves the pledge and then publishes the message. It passes every test, because no test stops the process between the two lines.

Production does. A deploy, an out-of-memory kill, or a broker that is restarting lands between the save and the publish, and the pledge exists with nobody told: no match, no receipt. Put the publish first and the same crash does the opposite: in the lesson’s example, the sponsor matches 5000 cents for a pledge the platform never recorded. The database’s transaction cannot help, because the broker is not in it.

The brief never asked the question: if the process dies between saving and announcing, which one survives, and who finishes the other?

02 / Name the shape

Write the pledge and its message together. Let a relay send the message.

A transactional outbox is a table in the service’s own database that holds messages waiting to go out. The request writes the business row and the outbox row in one transaction, so they exist together or not at all. A relay reads pending rows, publishes each one, and marks it sent. The request never talks to the broker, so a broker outage delays announcements instead of losing them.

The change and its message commit together. The relay publishes, then marks. Every message may arrive twice, so every receiver skips an id it has already handled.

Who owns what:

What each part of the outbox owns
PartOwnsPromises
The requestThe pledge and its outbox rowBoth rows in one transaction, or neither
The outbox tableMessages not yet sentA row stays pending until the broker has taken it
The relayGetting pending rows to the brokerIn order, retried until accepted, marked only after
Each receiverWhat it does with a messageA message id it has handled changes nothing

Words to put in a prompt or a review

Dual write
Writing to two systems one after the other and hoping nothing happens in between.
Outbox row
The message, with its id, stored in the same database and transaction as the change it announces.
Relay
The loop that publishes pending rows. Also called a message relay or outbox publisher.
Pending age
How long the oldest unsent row has waited. It is the one number to alert on.
At-least-once
Every message is delivered, some of them twice. It is what an outbox can promise.
Why not publish inside the database transaction?Two systems

Calling the broker between BEGIN and COMMIT does not make it part of the transaction. If the commit fails after the publish, the message is out and the pledge is not; the same ghost, a few lines later. It also holds database locks while waiting on the network. The outbox works because both writes go to one system that can commit them together.

03 / A crash between the writes

Same pledges, same crash. Which platform still tells the truth?

Each column runs one of the lesson’s platforms on the same steps, and audits the database against the broker: every recorded pledge announced, nothing announced that was never recorded. Watch four situations, then open Try it and crash them yourself.

Transactional outbox

A pledge, a crash, and the announcement

Save, then publish

Pledge reply: crashed

Database

  • pledge p-1

Broker

  • no messages

Matching fund: 0 cents

Save the message with the pledge

Pledge reply: crashed

Database

  • pledge p-1
  • outbox msg-1 · pending

Broker

  • no messages

Matching fund: 0 cents

01/ 04
Saved, then the process dies

Ada pledges 5000 cents, and the process dies after its first write.

Ada pledges 5,000 cents and the process dies straight after its first write. Saved alone, the pledge exists and nobody hears of it: no match, no receipt. Saved with its message, the relay finds the row and announces it.

Reduced motion: choose a scene to see its completed state.

Read this scene

Ada pledges 5,000 cents and the process dies straight after its first write. Saved alone, the pledge exists and nobody hears of it: no match, no receipt. Saved with its message, the relay finds the row and announces it.

Save, then publish: pledges p-1, broker messages none, matched 0 cents.

Save the message with the pledge: pledges p-1, broker messages none, matched 0 cents.

Watch restarts the story when you come back. Step through shows where each chapter ends. Try it starts a new platform whenever you change the design or press Reset.

04 / Read the shape

One transaction for two rows, and a relay that marks only what the broker took.

Basic form records a pledge. In the wild is the relay. At the call site shows who calls each. Notice that no line in the request mentions the broker.

Recording a pledge: the pledge row and its outbox row are written in one transaction, with the message id fixed there. The request never calls the broker, so a broker outage cannot fail it.

TypeScriptReading
platform.ts
/**
 * The pledge and its message are one transaction: both rows, or neither.
 * The message id is fixed here, so every later attempt to publish it
 * carries the same id. The request never talks to the broker.
 */
protected recordPledge(request: PledgeRequest): Reply {
	const pledge = this.pledgeOf(request);
	this.transaction((tables) => {
		tables.pledges.push(pledge);
		const id = `msg-${tables.outbox.length + 1}`;
		tables.outbox.push({ id, pledge: pledge.id, amount: pledge.amount, published: false });
	});
	this.crashAt(request, 'after-commit');
	return { status: 'recorded' };
}
GoAlongside
platform.go
// Record writes the pledge and its message in one transaction: both rows, or
// neither. The message id is fixed here, so every later attempt to publish it
// carries the same id. The request never talks to the broker.
func (p *OutboxPlatform) Record(r PledgeRequest) string {
	pl := pledgeOf(r)
	p.transaction(func(t *Tables) error {
		t.Pledges = append(t.Pledges, pl)
		id := "msg-" + itoa(len(t.Outbox)+1)
		t.Outbox = append(t.Outbox, OutboxRow{Message{id, pl.ID, pl.Amount}, false})
		return nil
	})
	return outcome("recorded", crashAt(r, "after-commit"))
}
The two dual writesSave, then publish; publish, then save

Both look correct, and both are, until the process dies between their two lines. The first loses the announcement; the second announces a pledge that does not exist.

platform.ts
/** Save, then publish: a crash in between records a pledge nobody hears about. */
export class SaveThenPublish extends Platform {
	protected recordPledge(request: PledgeRequest): Reply {
		const pledge = this.pledgeOf(request);
		this.transaction((tables) => tables.pledges.push(pledge));
		this.crashAt(request, 'after-save');
		const id = `msg-${this.broker.messages.length + 1}`;
		const ok = this.broker.publish({ id, pledge: pledge.id, amount: pledge.amount });
		return { status: ok ? 'recorded' : 'failed' };
	}
}

/** Publish, then save: a crash in between announces a pledge that was never recorded. */
export class PublishThenSave extends Platform {
	protected recordPledge(request: PledgeRequest): Reply {
		const pledge = this.pledgeOf(request);
		const id = `msg-${this.broker.messages.length + 1}`;
		if (!this.broker.publish({ id, pledge: pledge.id, amount: pledge.amount }))
			return { status: 'failed' };
		this.crashAt(request, 'after-publish');
		this.transaction((tables) => tables.pledges.push(pledge));
		return { status: 'recorded' };
	}
}
The behavior these examples promiseChecked by 18 shared scenarios
  • Save, then publish: a crash after the save, or a broker that is down, leaves a pledge that was never announced, and nothing tries again.
  • Publish, then save: a crash after the publish leaves an announcement, and a match, for a pledge that was never recorded.
  • The outbox: every recorded pledge is announced once the relay runs, whatever crashed and whenever the broker was down. A relay that stops after publishing sends the message again with the same id, and the matching fund matches it once.

Every expectation was generated by a separate model written from the contract in the examples’ README, not copied from either implementation, and it is kept beside the examples.

Reading the TypeScriptA transaction as a copy

transaction runs the work on a structuredClone of the tables and swaps it in only if the work finishes, so a crash inside it leaves nothing behind, the way a rolled-back transaction does. A crash is a thrown Crash, caught by record and reported as crashed.

Reading the GoThe same copy, and an error

Go copies the two slices, runs the work, and assigns the copy back only when the work returns no error. A crash is errCrash, returned by the step where the process dies. The three designs share a base struct through embedding.

Run it yourselfNo dependencies

Save the complete files at the paths in their banners. Then run node --experimental-strip-types run.ts (Node 22.18 or later), or go run . in the Go folder. Both print:

save-then-publish · crash between the writes: pledges [p-1], messages [], matched 0 -> never-announced:p-1
publish-then-save · crash between the writes: pledges [], messages [msg-1], matched 5000 -> ghost:p-1
outbox · crash after the commit: pledges [p-1], messages [msg-1], matched 5000 -> consistent
save-then-publish · broker down, then up: pledges [p-1], messages [], matched 0 -> never-announced:p-1
outbox · broker down, then up: pledges [p-1], messages [msg-1], matched 5000 -> consistent
outbox · relay stops, runs again: pledges [p-1], messages [msg-1 msg-1], matched 5000 -> consistent

05 / Review the agent’s diff

“Marking first fixes it.”

A donor got two receipts, and the agent found why. Read what its fix does to a crash.

The agent’s pull request

“The receipts service emailed a donor twice after a deploy. I traced it to the relay: it publishes and then marks the row, so a restart between the two sends it again. Marking first fixes it. All relay tests pass.”

// relay.ts
			for (const row of await outbox.pending()) {
			(removed)   await broker.publish(row.id, row.payload);
			(removed)   await outbox.markPublished(row.id);
			(added)   await outbox.markPublished(row.id); // mark first, so a crash never sends twice
			(added)   await broker.publish(row.id, row.payload);
			}
			
The relay tests never stop the process mid-loop. What do you do with this change?

06 / How it fails

Every failure here is a crash in the gap: the question is what was written down.

Each row except the last two is a shared scenario the tests run.

Failure modes of saving a pledge and announcing it
What happensSave, then publishPublish, then saveThe outbox
The process dies after its first writeA pledge nobody hears about.A match for a pledge that does not exist.Both rows committed; the relay announces it.
The broker is down for a minuteThe pledge is saved, the publish fails, and nothing retries.The request fails; the donor tries again.The request succeeds; the row waits for the broker.
The relay stops after a publish, before the markNo relay.The message goes twice with one id; the fund matches once.
A receiver does not check idsRarely matters: each message is sent at most once.A second receipt, or a second match. Not modeled in the example.
The relay is stuckNot applicable.Pending rows pile up and announcements stop, silently unless someone watches pending age. Not modeled in the example.

The receiving side has its own lessons: Idempotency and at-least-once explains why a message must be safe to receive twice, and Receiving webhooks keeps a receipt table for exactly that.

07 / Is it worth it?

A table and a relay, against two lines that are almost always fine.

The shared four changes, against a dual write
ChangeDual writeThe outbox
A second entry point: pledges imported from a partner’s CSVThe import must remember to publish, and has its own gap.The import writes the same two rows; the relay does the rest.
A new receiver: a leaderboardNo change.No change.
A new rule: a canceled pledge is announced tooA second dual write with the same gap.A second message type in the same transaction as the cancel.
The broker is replacedEvery call site changes.The relay changes; requests do not.

The costs are real: a table that grows until something deletes published rows, a relay to run and watch, announcements that lag by one relay turn, and receivers that must handle a copy. When a lost or phantom message costs nothing, say a cache hint or an analytics ping, the dual write is fine. When it is money or a promise to a person, it is not.

Measure before and after:

  • Records with no message, and messages with no record, from a daily reconciliation. With a dual write, this is the number that is not zero.
  • Age of the oldest pending row, which is how late announcements are running, and duplicate messages per day seen by receivers.
  • Outbox table size, against your cleanup policy.

This lesson did not measure a real platform, and gives no numbers.

08 / Ask for it

One brief, two prompts.

Two agents running Claude Sonnet each got the brief from section 01. One prompt added an Outbox block: the pledge and its message in one transaction, a relay that publishes in order and marks after the broker accepts, a message id fixed when the row is written, and a broker that never slows a pledge. A script ran both builds against its own broker, which it took down, slowed, and made hold a message without answering, and it killed the servers along the way.

What the checker found, run 2026-09-23
QuestionPlain promptOutbox prompt
Ten pledges, one after another10 of 10 announced, out of order; replies in 4 ms or less10 of 10 announced, in order; replies in 2 ms or less
The broker down for 20 secondsEvery pledge answered 201; 5 of 5 announced 1010 ms after it came back, out of orderEvery pledge answered 201; 5 of 5 announced 202 ms after it came back, in order
A broker that takes 8 seconds to answerReplies in 2 ms or less; 3 of 3 announced, 7 messages sent for 3 pledgesReplies in 2 ms or less; 3 of 3 announced, 3 messages sent for 3 pledges
Server killed while the broker holds a publish1 of 1 announced; sent again after the restart with the same id (1 copy)1 of 1 announced; sent again after the restart with the same id (1 copy)
Server killed with five pledges pending5 of 5 announced5 of 5 announced
Twenty pledges at once20 of 20 announced; one id per pledge20 of 20 announced; one id per pledge
Its own tests14 of 14 pass19 of 19 pass

Both agents built an outbox. The brief said “a pledge must never be announced unless it was recorded, and a recorded pledge must be announced”, and that sentence was enough: the plain agent wrote the pledge and its message in one transaction, marked a row only after a 202, and resumed after a restart. No pledge was lost and none was announced without being recorded, in either build, under any question.

The difference was order, and the brief never mentioned it. The plain build gives each outbox row a random UUID and publishes pending rows sorted by that id, so pledges reach the broker in an order nobody chose. For a matching fund that adds up totals, that is harmless. For a receiver that sees “pledge canceled” before “pledge recorded”, it is not.

src/db.ts · plain prompt
  const messageId = randomUUID();
…
    .prepare('SELECT id, pledge_id, payload FROM outbox WHERE published = 0 ORDER BY id LIMIT ?')

The outbox build keeps a separate sequence number for the order and the message id for deduplication, and publishes one row at a time, oldest first.

lib/db.ts · outbox prompt
  // seq is the durable publish order: rows are relayed in ascending seq
  // order. id is the message id that goes out to the broker; it is fixed
  // the moment the row is written, so republishing the same row (after a
  // crash, or a lost 202) always carries the same id and downstream
  // consumers can dedupe on it.
  db.exec(`
    CREATE TABLE IF NOT EXISTS outbox (
      seq INTEGER PRIMARY KEY AUTOINCREMENT,
      id TEXT NOT NULL UNIQUE,
      topic TEXT NOT NULL,

The missing line is one sentence: publish pending rows in the order they were written, by a sequence, not by the message id. The rest of the outbox block restated what the plain agent already did. A brief that states the guarantee gets most of the mechanism; it does not get the properties nobody wrote down.

The slow broker showed the other unstated number. The plain build gives up on a publish after 5 seconds and tries again, so an 8-second broker got several copies of each message; the outbox build waits 10 seconds and sent each once. The copies carried the same id, so nothing downstream was wrong, but the timeout against the broker’s real latency is a number to ask for.

How the runs were made and checkedTwo builds, recorded as written
  • Both agents were launched at the same time from empty folders; neither was told about the other, the lesson, or the checker.
  • Both builds are kept byte for byte with checksums. For every question the checker restores a build into a fresh folder with its own database, starts the server, and runs its own broker.
  • The checker ran twice; the first covered the plain build alone while the outbox agent was still working, and gave the same answers apart from timings.
  • Both agents wrote a few scratch files in /tmp during their own checks, against the prompt, and deleted them. Neither stopped a process by name or pattern.
  • One run of each prompt is a sample, not a measurement of the model.

09 / Hold it there

An outbox breaks when someone adds a publish outside it. Three checks notice.

  1. Only the relay imports the broker client

    The moment a request handler calls the broker directly, the dual write is back. An import rule that allows the broker client in the relay and nowhere else keeps it out, the way Enforcement layer enforces import rules.

  2. Tests that stop the process between the steps

    A test that records a pledge and waits proves nothing about the gap. The shared scenarios crash after each write, take the broker down, and stop the relay between its publish and its mark; a check that the pledge and its outbox row are written by one transaction catches the rest.

  3. Alert on pending age, and reconcile

    Alert when the oldest pending row is older than a relay turn or two. Once a day, compare pledges with the messages the matching fund received; with an outbox the difference should be the rows still pending, and nothing else.

There is no frontend version of this lesson: the outbox sits between a service’s database and its broker, which a browser never sees. The browser’s own version, a change and its unsent request kept together in local storage until the server takes them, is in Local-first and sync.

10 / Make the call

If the message matters, it goes in the transaction.

Use an outbox whenever a change must be followed by a message that another system acts on: money, emails to people, stock, access. Keep the dual write for messages whose loss nobody would notice. Reopen the design when the database already publishes its own changes; then Change data capture can play the relay.

Take it with you

Explain it without saying “outbox”: “When we take a pledge, we write the pledge and a note saying ‘tell the sponsor about this’ on the same page, in one go. Someone reads the notes and passes each one on, ticking it off after. If they are interrupted, they pass it on again, and the sponsor ignores a note it has already seen.” Then find a place in your own code that saves something and then sends a message, and ask what happens if the process dies on the line between them.

Paste into your next prompt, and fill in the blanks

When [a pledge] is recorded, [the matching fund and receipts service] must hear about it through [the broker].
Write the [pledge] and its message to an outbox table in the same database transaction; the request never calls the broker.
Give each message its id when the row is written, and publish it with that id every time.
A relay in the same service publishes pending rows in the order they were written, by a sequence number, marks each one only after the broker accepts it, and keeps trying while the broker is down; a restart resumes from the rows still pending.
Receivers may get a message twice: they skip an id they have already handled.
Alert when the oldest pending row is older than [one minute], and delete published rows after [seven days].
Connections to follow nextRelated lessons

Take the outbox into your editor. Run two relays at once, and make them share the work without publishing a row twice or out of order.

Back to architecture →