← Architecture
From modular monolith to services Between two commits

Consistency without a shared transaction

Write it down first. Finish it on restart.

Today your checkout holds the stock and writes the order in one transaction, because both tables sit in one database. It is the quiet reason the two never disagree. Give one module its own database, and that guarantee disappears without a single line of checkout code changing.

TypeScriptGoOne checkout, three designs, a crash at every step, two recorded builds.

01 / The prompt

“Give Inventory its own database. It becomes a service next quarter.”

It is a sensible step on the way to a service, and an agent will do it cleanly: a second database file, Inventory’s tables moved into it, every test still green. The checkout code does not even change. It still holds a mug, charges the card, and writes the order, in that order.

What changed is what happens between those steps. Before, the hold and the order were one transaction; a crash rolled both back. Now each is its own commit. Stop the process after the charge and the mug stays held for an order that was never written, and the card provider has Ada’s money. In the lesson’s shop that one crash leaves 2 broken promises, and pressing Pay again takes her money a second time: 2 charges for one mug.

The ticket never asked the question: when a checkout stops halfway, who finishes it, and how do they know it was there?

02 / Name the shape

Write the checkout down before you start it, and let startup finish what a crash left.

A transaction makes several writes land together, but only inside one database. Once a business action crosses two databases, or a provider, it is a series of commits, and the process can stop between any two. Consistency then comes from three things you build: a record of the action written first, a key that makes every step safe to repeat, and a pass that reads the record and finishes or undoes what is half done.

Record the checkout as pending in Orders’ own database before touching anything else. Name the hold and the charge after its id. On startup, ask the provider about every pending order, then place it or release its stock.

Who owns each fact, once Inventory has its own file:

Who owns each fact in the checkout, and where it is committed
FactOwner and fileWhat the workflow relies on
That a checkout started, and how far it gotOrders, orders.dbThe pending order, written before the first other step. Its id is the key.
Stock, and which units are heldInventory, inventory.dbA hold named after the order id, so asking twice returns the same hold. Release and settle do nothing the second time.
Whether the card was chargedThe card provider, its own recordsA charge sent with the order id as its key, and a lookup by that key. Only the provider knows what happened after a lost reply.
Whether the order is placedOrders, orders.dbChanged only after the provider has answered, by the checkout or by recovery.

Words to put in a prompt or a review

Commit point
The moment one database makes a write permanent. A checkout across two has several.
Invariant
What must be true when nothing is in progress: every hold and charge has an order.
Intent record
The pending order written first, so a restart knows a checkout existed.
Idempotency key
An id sent with a step so that repeating it returns the first result, not a second one.
Recovery pass
Code that runs on startup and moves every pending record to placed or failed.
Unknown outcome
A call with no reply. It might have worked; only the other side can say.
Why not one transaction across both?Two-phase commit, and attached files

Some databases can coordinate one commit across several: two-phase commit, or SQLite’s ATTACH. They cover databases, not a card provider, and they tie the two modules to one coordinator, which is the coupling the split was meant to remove. SQLite’s own documentation is careful about even the local case: “If the main database is ":memory:" or if the journal_mode is WAL, then transactions continue to be atomic within each individual database file. But if the host computer crashes in the middle of a COMMIT where two or more database files are updated, some of those files might get the changes where others might not.” (SQLite, ATTACH DATABASE, checked 23 September 2026)

When the steps are many and undoing one is a business action of its own, such as a refund or a canceled booking, the same idea grows into a saga with compensations. Sagas and compensation covers that on a trip booking.

03 / Stop the checkout

Same crash, three designs. What does each leave on disk?

Ada buys one mug. Each chapter runs the lesson’s own checkout, stops the process at a step, and starts the shop again from its files. The boxes are what each file and the provider hold after every step. Watch, then open Try it and pull the plug wherever you like.

Consistency

Stop the checkout between two commits

Process: Running

shop.db

orders_orders

  • none

inventory_holds · mug stock 12

  • none

In the open transaction: hold hold-1

provider

charges

  • none
01/ 04
Today: one file, one transaction

Inventory holds a mug inside the transaction, not committed yet.

Both modules’ tables share shop.db, so the hold and the order commit together or not at all. The crash rolls both back. The card was never inside that transaction.

Reduced motion: choose a scene to see its completed state.

Read this scene

Both modules’ tables share shop.db, so the hold and the order commit together or not at all. The crash rolls both back. The card was never inside that transaction.

Inventory holds a mug inside the transaction, not committed yet.

Orders: none. Holds: none. Charges: none.

Watch restarts the story when you come back. Step through shows where each chapter ends. Try it starts a fresh shop every time you run it.

04 / Read the shape

A record first, keys on every step, and a pass that runs before the first request.

Basic form is the checkout written as commits in an order you chose. In the wild is the recovery pass that makes that order safe. At the call site is where the shop opens its files and runs it. Notice what gets written before the provider is called, and what recovery asks the provider rather than guesses.

The checkout after the split. Orders writes the order as pending before anything else, and its id names the hold and keys the charge. Each step commits on its own.

TypeScriptReading
checkout.ts
/**
 * After the split. Orders records the checkout before anything else, and the
 * order id is the key every later step uses: Inventory names the hold after it,
 * and the provider gets it as the idempotency key. Each step commits on its own;
 * a crash between two of them leaves a pending order that says how far it got.
 */
function workflow(shop: Shop, request: Request): Reply {
	const seen = shop.orders.read().orders_orders.find((order) => order.key === request.key);
	if (seen) return replyFor(seen); // the same checkout, sent again
	const total = prices[request.sku] * request.qty;

	const orderId = shop.orders.transaction((tables) => {
		const id = nextOrderId(tables);
		tables.orders_orders.push(order(id, request, total, id, null, 'pending'));
		return id;
	});
	note(shop, { what: 'record', where: 'orders.db', committed: true, id: orderId });
	at(shop, 'recorded');

	const held = shop.inventory.transaction((tables) =>
		hold(tables, request.sku, request.qty, orderId)
	);
	if (held.status === 'out-of-stock') {
		shop.orders.transaction((tables) => finish(tables, orderId, 'failed', 'out-of-stock'));
		note(shop, { what: 'fail', where: 'orders.db', committed: true, id: orderId });
		return { status: 'rejected', reason: 'out-of-stock' };
	}
	note(shop, { what: 'hold', where: 'inventory.db', committed: true, id: orderId });
	at(shop, 'reserved');

	let charge: Charge;
	try {
		charge = shop.provider.charge(total, request.card, orderId);
	} catch (error) {
		if (!(error instanceof ReplyLost)) throw error;
		// Unknown is not declined. The order stays pending, and recover() asks the provider.
		note(shop, { what: 'reply-lost', where: 'provider', committed: true, id: orderId });
		return { status: 'pending', orderId };
	}
	if (charge.status !== 'succeeded') {
		shop.inventory.transaction((tables) => release(tables, orderId));
		shop.orders.transaction((tables) => finish(tables, orderId, 'failed', 'declined'));
		note(shop, { what: 'fail', where: 'orders.db', committed: true, id: orderId });
		return { status: 'rejected', reason: 'declined' };
	}
	note(shop, { what: 'charge', where: 'provider', committed: true, id: charge.id });
	at(shop, 'charged');

	shop.orders.transaction((tables) => finish(tables, orderId, 'placed', null, charge.id));
	note(shop, { what: 'place', where: 'orders.db', committed: true, id: orderId });
	at(shop, 'placed');

	shop.inventory.transaction((tables) => settle(tables, orderId));
	note(shop, { what: 'settle', where: 'inventory.db', committed: true, id: orderId });
	return { status: 'placed', orderId, total };
}
GoAlongside
checkout.go
// workflow is after the split. Orders records the checkout before anything
// else, and the order id is the key every later step uses: Inventory names the
// hold after it, and the provider gets it as the idempotency key. Each step
// commits on its own; a crash between two of them leaves a pending order that
// says how far it got.
func (s *Shop) workflow(r Request) (Reply, error) {
	for _, o := range s.orders.Read().Orders {
		if o.Key == r.Key {
			return replyFor(o), nil // the same checkout, sent again
		}
	}
	total := prices[r.Sku] * r.Qty

	var orderID string
	s.orders.Transaction(func(t *OrdersTables) error {
		orderID = t.NextID()
		t.Orders = append(t.Orders, newOrder(orderID, r, total, orderID, "", "pending"))
		return nil
	})
	s.note("record", "orders.db", orderID, true)
	if err := s.at("recorded"); err != nil {
		return Reply{}, err
	}

	var held string
	s.inventory.Transaction(func(t *InventoryTables) error {
		held = t.HoldStock(r.Sku, r.Qty, orderID)
		return nil
	})
	if held == "" {
		s.finish(orderID, "failed", "out-of-stock", "")
		s.note("fail", "orders.db", orderID, true)
		return Reply{Status: "rejected", Reason: "out-of-stock"}, nil
	}
	s.note("hold", "inventory.db", orderID, true)
	if err := s.at("reserved"); err != nil {
		return Reply{}, err
	}

	charge, err := s.Provider.Charge(total, r.Card, orderID)
	if errors.Is(err, ErrReplyLost) {
		// Unknown is not declined. The order stays pending, and recover asks the provider.
		s.note("reply-lost", "provider", orderID, true)
		return Reply{Status: "pending", OrderID: orderID}, nil
	}
	if charge.Status != "succeeded" {
		s.inventory.Transaction(func(t *InventoryTables) error { t.Release(orderID); return nil })
		s.finish(orderID, "failed", "declined", "")
		s.note("fail", "orders.db", orderID, true)
		return Reply{Status: "rejected", Reason: "declined"}, nil
	}
	s.note("charge", "provider", charge.ID, true)
	if err := s.at("charged"); err != nil {
		return Reply{}, err
	}

	s.finish(orderID, "placed", "", charge.ID)
	s.note("place", "orders.db", orderID, true)
	if err := s.at("placed"); err != nil {
		return Reply{}, err
	}

	s.inventory.Transaction(func(t *InventoryTables) error { t.Settle(orderID); return nil })
	s.note("settle", "inventory.db", orderID, true)
	return Reply{Status: "placed", OrderID: orderID, Total: total}, nil
}
The checkout today, in one fileOne transaction, and the card outside it

While both modules share shop.db, the hold, the order, and the settle happen inside one transaction. A crash anywhere inside it writes nothing. The charge is the one step that is not undone, which is why even today a crash after it leaves a charge with no order.

checkout.ts
/** Today: both modules' tables live in one file, so one transaction covers the stock and the order. */
function oneFile(shop: Shop, request: Request): Reply {
	const total = prices[request.sku] * request.qty;
	const step = (what: Step['what'], id: string) =>
		note(shop, { what, where: 'shop.db', committed: false, id });
	try {
		return shop.together!.transaction((tables) => {
			const orderId = nextOrderId(tables);
			const held = hold(tables, request.sku, request.qty);
			if (held.status === 'out-of-stock') throw new Rejected('out-of-stock');
			step('hold', held.holdId);
			at(shop, 'reserved');
			// The card provider is not in the transaction. Nothing can roll this back.
			const charge = chargeOrNull(shop, total, request.card);
			if (charge?.status !== 'succeeded') throw new Rejected('declined');
			note(shop, { what: 'charge', where: 'provider', committed: true, id: charge.id });
			at(shop, 'charged');
			tables.orders_orders.push(order(orderId, request, total, held.holdId, charge.id));
			step('place', orderId);
			at(shop, 'placed');
			settle(tables, held.holdId);
			step('settle', held.holdId);
			return { status: 'placed', orderId, total };
		});
	} catch (error) {
		if (error instanceof Rejected) return { status: 'rejected', reason: error.reason };
		throw error;
	}
}
The same steps after the splitWhat the ticket produces by default

The code a straight move produces: the same order of steps, each now committed on its own. It passes every happy-path test the one-file version passes.

checkout.ts
/** The same steps after Inventory moves to its own file: every write is its own commit. */
function twoFiles(shop: Shop, request: Request): Reply {
	const total = prices[request.sku] * request.qty;
	const orderId = nextOrderId(shop.orders.read());
	const held = shop.inventory.transaction((tables) => hold(tables, request.sku, request.qty));
	if (held.status === 'out-of-stock') return { status: 'rejected', reason: 'out-of-stock' };
	note(shop, { what: 'hold', where: 'inventory.db', committed: true, id: held.holdId });
	at(shop, 'reserved');
	const charge = chargeOrNull(shop, total, request.card);
	if (charge?.status !== 'succeeded') {
		shop.inventory.transaction((tables) => release(tables, held.holdId));
		note(shop, { what: 'release', where: 'inventory.db', committed: true, id: held.holdId });
		return { status: 'rejected', reason: 'declined' };
	}
	note(shop, { what: 'charge', where: 'provider', committed: true, id: charge.id });
	at(shop, 'charged');
	shop.orders.transaction((tables) => {
		tables.orders_orders.push(order(orderId, request, total, held.holdId, charge.id));
	});
	note(shop, { what: 'place', where: 'orders.db', committed: true, id: orderId });
	at(shop, 'placed');
	shop.inventory.transaction((tables) => settle(tables, held.holdId));
	note(shop, { what: 'settle', where: 'inventory.db', committed: true, id: held.holdId });
	return { status: 'placed', orderId, total };
}
The auditThe invariant, as code over both files and the provider

Every scenario ends with this check. A held mug is fine while its order is pending, and a broken promise otherwise; a successful charge must belong to a placed order, or to a pending one under its own key.

audit.ts
import type { Shop } from './checkout.ts';
import type { Hold } from './inventory.ts';
import type { Charge } from './provider.ts';

// The invariant, as a check over what the disks and the provider hold:
// every hold ends sold with a placed order or released, and every successful
// charge belongs to a placed order. A pending order may still hold stock, or
// have a charge made under its id: that is a checkout in progress, not a broken one.

export type Snapshot = {
	stock: Record<string, number>;
	orders: {
		id: string;
		state: string;
		reason: string | null;
		holdId: string;
		chargeId: string | null;
	}[];
	holds: Hold[];
	charges: Charge[];
	violations: string[];
};

export function audit(shop: Shop): Snapshot {
	const inventory = shop.inventory.read();
	const orders = shop.orders
		.read()
		.orders_orders.map(({ id, state, reason, holdId, chargeId }) => ({
			id,
			state,
			reason,
			holdId,
			chargeId
		}));
	const charges = shop.provider.charges.map((charge) => ({ ...charge }));
	const violations: string[] = [];
	for (const hold of inventory.inventory_holds) {
		if (hold.state !== 'held') continue;
		const owner = orders.find((order) => order.holdId === hold.id);
		if (owner?.state === 'pending') continue;
		violations.push(
			owner?.state === 'placed' ? `hold-not-settled:${owner.id}` : `held-without-order:${hold.id}`
		);
	}
	for (const charge of charges) {
		if (charge.status !== 'succeeded') continue;
		if (orders.some((order) => order.state === 'placed' && order.chargeId === charge.id)) continue;
		if (orders.some((order) => order.state === 'pending' && order.id === charge.key)) continue;
		violations.push(`charge-without-order:${charge.id}`);
	}
	return {
		stock: inventory.inventory_stock,
		orders,
		holds: inventory.inventory_holds,
		charges,
		violations
	};
}
The behavior these examples promiseChecked by 38 shared scenarios
  • A paid order, a declined card, too little stock, and a sold-out product end the same way in all three designs.
  • Stopped after any step and restarted, the workflow always ends consistent: the order placed when the provider has a charge under its id, failed with the mug back in stock when it does not.
  • The same checkout sent again returns the first answer and charges nothing more in the workflow, and places a second order in the other two.
  • A lost reply from the provider is a decline to the first two designs, with the money kept; the workflow answers pending and finishes the order on restart.

Every expectation was generated by a separate model written from the contract in the examples’ README, not copied from either implementation, and it is kept beside the examples.

Reading the TypeScriptA file per database, a copy per transaction

A database here is one JSON file. transaction(work) reads the committed copy, runs work on it, and writes the file once; if work throws, nothing is written. A crash is a thrown Crash: the shop object is dropped, and a new one opens the same files. The provider is passed to every restart, because its records belong to another company and survive our crashes.

Reading the GoErrors instead of exceptions

The same files, as a generic Database[T] whose Transaction writes only when its work returns no error. A crash is a Crash error returned up the stack. FileDisk commits by writing a temporary file and renaming it over the old one, so a reader never sees half a file.

Run it yourselfNo dependencies

Save the complete files at the paths in their banners. Then run node --experimental-strip-types run.ts (Node 22.18 or later), or go run . in the Go folder. Both write real files in a temporary folder and print:

one file, crash after charged: mug stock 12, orders none, charges ch-1 succeeded -> charge-without-order:ch-1
two files, crash after charged: mug stock 11, orders none, charges ch-1 succeeded -> held-without-order:hold-1, charge-without-order:ch-1
workflow, crash after charged: mug stock 11, orders order-1 placed, charges ch-1 succeeded -> consistent
workflow, crash after reserved: mug stock 12, orders order-1 failed, charges none -> consistent
workflow, reply lost: mug stock 11, orders order-1 placed, charges ch-1 succeeded -> consistent

05 / Review the agent’s diff

“Checkout now undoes the hold if anything fails. All tests pass.”

The test it added is real, and it passes. Read the diff for the failure it cannot see.

The agent’s pull request

“Checkout now undoes the stock hold if payment or saving the order fails. I added a test that makes orders.place throw, and the mug is back in stock afterwards. All tests pass.”

// orders/checkout.ts
			const held = inventory.hold(sku, qty);
			(added) try {
			  const charge = provider.charge(total, card);
			  orders.place(orderId, held.holdId, charge.id);
			  inventory.settle(held.holdId);
			(added) } catch (error) {
			(added)   // Undo the hold if anything after it fails.
			(added)   inventory.release(held.holdId);
			(added)   throw error;
			(added) }
			
Inventory moves to its own file next week. What do you do with this change?

06 / How it fails

The failures that matter happen between two commits, and nobody sees them at the time.

Each row is a shared scenario the tests run. “Same steps” is the checkout moved across as it was; “workflow” records first and recovers on start.

Failure modes of one checkout across two databases and a card provider
What goes wrongWhat Ada seesSame steps, after a restartWorkflow, after a restart
Stops after the mug is heldAn error, or a spinner that never endsThe mug stays held for no order, forever.Nothing was charged, so recovery releases the mug and fails the order as interrupted.
Stops after the card is chargedThe sameCharged with no order, and the mug held. Nothing will ever notice.Recovery finds the charge under the order’s id and places the order.
Stops after the order is writtenThe sameThe order exists, and its hold is never settled; a cleanup job that frees old holds would sell the mug twice.Recovery settles the hold of every placed order.
The provider’s reply is lostSame steps: “declined”. Workflow: “we are confirming your payment”.Treated as a decline: the mug is released and the charge kept.The order stays pending, and recovery places it from the provider’s record.
Ada presses Pay again after the crashSame steps: a success, for the second chargeA second hold and a second charge; the first ones stay broken.The same order id comes back. Nothing is charged twice.
The crash happens today, in one fileAn errorThe transaction rolls back the hold and the order together, and the charge stays: the card was never in the transaction. The split did not create this gap, only made it wider.
The provider is slowA long waitThe mug is held while Ada waits.The same, with a pending order that says why. Not modeled in the example; set a timeout and treat it as an unknown outcome.

Retries, timeouts, and keys each have a lesson of their own: Idempotency and at-least-once, Timeouts, deadlines, and races, and Transactions and atomicity.

07 / Is it worth it?

A record and a recovery pass cost a table and a startup step. What do they buy?

The shared four changes, against the same steps and the workflow
ChangeSame stepsWorkflow
A second entry point: the mobile app’s checkoutEach client needs its own story for a retry after a timeout, and gets it wrong in its own way.Both send a cart key; the server answers a repeat with the first result.
A new card providerSwap the call.Swap the call, and check the new provider can look a charge up by your key. Recovery depends on it.
A new rule: holds expire after 15 minutesAn expiry job, which can free the mug of an order that was placed but never settled.The job skips holds whose order is pending or placed, because the record says which is which.
Inventory moves to another team, as a serviceThe checkout has to be rethought first.No change: Inventory was already a separate commit, reached with a key.

The cost is real: a pending state that every screen and report has to understand, a recovery pass that has to be correct, and a provider that has to support keys. For a checkout that never leaves one database and never calls out mid-transaction, one transaction is simpler and enough.

Decide what “consistent” means before the split, then measure it, before and after:

  • Audit findings per day: holds with no order, charges with no placed order, and placed orders with an unsettled hold, from a query like the audit above run against production data and the provider’s own charge list. Run it on today’s one-file shop first; a charge with no order is already possible there.
  • Pending orders older than five minutes. After a restart the target is zero; a number that grows means recovery is failing.
  • Charges per placed order, which should be exactly one; above one means a retry charged twice.
  • Checkout time at the 95th percentile, since the record adds a write.

This lesson did not measure a real shop, and gives no numbers.

08 / Ask for it

One starting point, one ticket, two prompts.

Two agents running Claude Sonnet each got a copy of the shop as the Module-owned data run left it, and the same ticket: Inventory’s tables in their own database file, and card payments through a provider over HTTP, whose API the ticket describes, idempotency keys included. One prompt added a Consistency block: the invariant, a checkout recorded first under a key every step uses, unknown outcomes looked up by key, recovery on startup, and tests that crash after each step. A script ran both builds against its own payment provider, held a charge request open, stopped the server, and started it again on the same files.

What the checker found, run 2026-09-23
QuestionPlain promptConsistency prompt
Five orders, two of them refused201, 201, 402, 409, 201; every charge sent with a key201, 201, 402, 409, 201; every charge sent with a key
Stopped while the provider held the request, before charging; then restartedNo order, mug stock 11: one mug held for nothingThe order placed and charged once (mug stock 11)
Stopped after the provider charged, before it answered; then restartedNo order, mug stock 11: one mug held for nothing, and 1 charge kept with no refundThe order placed and charged once (mug stock 11)
A second restartNothing changesNothing changes
The provider charges and never answers201 after 5 s; one order, one charge201 after 5 s; one order, one charge
The provider is unreachable402 “declined” in 11 ms; the mug back in stockNo reply within 60 s; once the provider was back, the order was placed and charged
Its own tests67 of 67 pass82 of 82 pass
Crash points written downNoCONSISTENCY.md

Until the process stops, the plain build is careful. It sends every charge with a key and retries a lost reply with the same one, so the provider hands back the first charge instead of making a second. But the key is made fresh inside the call and lives only in memory, and nothing records that a checkout started. Stop the server while the provider holds the request, and after the restart the mug is still held for an order that does not exist; stop it after the charge, and Ada has paid for it too. Nothing in the build will ever look.

orders/index.ts · plain prompt
+  const paymentResult = await charge({ card: input.card, amountCents: total });
   if (!paymentResult.approved) {
     releaseStock(reserveResult.reservationId);
     return { status: "declined" };
payments/index.ts · plain prompt
+export async function charge(input: ChargeInput): Promise<ChargeResult> {
+  const idempotencyKey = randomUUID();

The consistency build wrote each checkout down first, keyed the reservation and the charge by it, and runs the same function for a new checkout and an interrupted one. Both crashes end with the order placed. Read what that means for the first one: the server stopped before the card was charged, Ada never got an answer, and after the restart her card was charged and the order placed. That keeps the invariant, and it is a product decision the prompt did not make. The lesson’s own example undoes such a checkout instead.

The same code never gives up on a charge whose outcome is unknown, so with the provider down, the request got no answer within a minute. The prompt said an unknown outcome is not a decline, and the agent followed it to the letter. It never said how long a customer should wait.

orders/index.ts · consistency prompt
+async function resolveCharge(idempotencyKey: string, amountCents: number, card: string): Promise<ResolvedCharge> {
+  let attempt = 0;
+  for (;;) {
+    const result = await requestCharge(idempotencyKey, amountCents, card);
+    if (result.outcome === "succeeded") {
+      return { outcome: "succeeded", chargeId: result.chargeId };
+    }
+    if (result.outcome === "declined") {
+      return { outcome: "declined" };
+    }
+
+    const found = await lookupCharge(idempotencyKey);
+    if (found) {
+      return { outcome: "succeeded", chargeId: found.chargeId };
+    }
+
+    attempt += 1;
+    const delay = Math.min(CHARGE_BASE_DELAY_MS * 2 ** attempt, CHARGE_MAX_DELAY_MS);
+    await sleep(delay);
+  }
server.ts · consistency prompt
+void recoverCheckouts().catch((err) => {
+  console.error("checkout recovery failed:", err);
+});

So the missing prompt lines are about the product, not the plumbing: a checkout interrupted before the charge is undone, not finished, and a customer waits at most N seconds, then sees “we are confirming your payment”, while recovery keeps working in the background. The prompt in section 10 includes both.

How the runs were made and checkedTwo builds, recorded as written
  • The starting point is the Module-owned data lesson’s recorded ownership build, byte for byte, with its checksums checked before the runs. Both agents were launched at the same time; neither was told about the other, the lesson, or the checker.
  • Both builds are kept byte for byte with checksums and diffs. For every question the checker restores a build into a fresh folder with its own two database files, runs its own mail and payment providers, and stops the server with a kill signal while the payment provider holds the request, either before charging or after.
  • After each restart it waits 12 seconds, reads stock, Ada’s orders, and the provider’s charges and refunds, restarts once more, and places one more order. The checker ran twice; the first run covered the plain build alone and gave the same answers.
  • The plain agent created and deleted one scratch file in /tmp. The consistency agent ran its crash tests on two ports it was not given, and inspected another agent’s server with ps and lsof without stopping it. Neither stopped a process by name.
  • One run of each prompt is a sample, not a measurement of the model.

09 / Hold it there

A consistency rule that only lives in a design document is a hope. Three checks keep it.

  1. The provider’s own door: idempotency keys

    The workflow leans on the provider remembering a key. Stripe documents it this way: “Stripe’s idempotency works by saving the resulting status code and body of the first request made for any given idempotency key, regardless of whether it succeeds or fails. Subsequent requests with the same key return the same result, including 500 errors.” Keys can be removed after they are at least 24 hours old, so recovery has to run well inside a day. (Stripe, Idempotent requests, checked 23 September 2026.) Check your provider’s documentation for the same two facts before you rely on them.

  2. Tests that crash at every step

    The shared scenarios stop the checkout after each step, restart, and run the audit; any violation fails the build. When a new step is added to checkout, the list of crash points grows with it. A rule that only Orders’ checkout may call the provider, enforced the way Enforcement layer runs import rules, keeps a second caller from charging without a key.

  3. The audit, against production

    Run the audit on a schedule against both databases and the provider’s list of charges, and alert on anything it finds and on pending orders past their age limit. This is the check on what actually happened, including crashes no test imagined.

Build UIs?The pay button is where a retry starts, and where “pending” has to be shown.

Where it already is in your components

Every pay button that disables itself while a request is in flight is guarding against the same double charge, one click at a time. It does nothing for a retry after a timeout, a refresh, or a second tab. Those are covered only when the page sends the same key each time and the server answers a repeat with the first result.

When you have to own it

Once the server can answer “pending”, the page has to say so honestly and ask again, rather than show a failure for what may be a success. You own the key: made once per checkout, kept across retries, and sent with every attempt.

A pay button that makes one key per checkout and sends it with every retry after a timeout.

ReactAlready in your code
PayButton.tsx
import { useRef, useState } from 'react';

type Reply =
	| { status: 'placed'; orderId: string }
	| { status: 'pending'; orderId: string }
	| { status: 'rejected'; reason: string };

// One key per checkout, made once. A double click, or a retry after a timeout,
// sends the same key, so the server finds the checkout it already started.
export function PayButton({ cartId, onDone }: { cartId: string; onDone: (reply: Reply) => void }) {
	const key = useRef(`${cartId}:${crypto.randomUUID()}`);
	const [state, setState] = useState<'idle' | 'sending' | 'retrying'>('idle');

	async function pay(attempt = 1): Promise<void> {
		setState(attempt === 1 ? 'sending' : 'retrying');
		try {
			const response = await fetch('/api/checkout', {
				method: 'POST',
				headers: { 'Content-Type': 'application/json', 'Idempotency-Key': key.current },
				body: JSON.stringify({ cartId }),
				signal: AbortSignal.timeout(10_000)
			});
			onDone((await response.json()) as Reply);
			setState('idle');
		} catch {
			// No reply is not a failure: the charge may have gone through. Ask again, same key.
			if (attempt < 3) return pay(attempt + 1);
			onDone({ status: 'pending', orderId: '' });
			setState('idle');
		}
	}

	return (
		<button type="button" disabled={state !== 'idle'} onClick={() => pay()}>
			{state === 'idle' ? 'Pay' : state === 'sending' ? 'Paying…' : 'Checking your payment…'}
		</button>
	);
}

10 / Make the call

Keep one transaction while you can. The day a write leaves it, write the checkout down.

While every table the checkout writes lives in one database, a transaction is the simplest correct answer; keep it, and keep the card call outside it. The day one write moves to another database or service, give the action a record written first, a key on every step, and a recovery pass that runs on startup. Reopen the decision when the steps multiply or an undo becomes a business action of its own; that is a saga.

Take it with you

Explain it without saying “consistency”: “Before we touch anything, we write down what we are about to do. If we crash, the first thing we do on the way back up is read that note and either finish the job or put everything back.” Then find the action in your own code that writes to two stores, and ask what a crash between them leaves behind.

Paste into your next prompt, and fill in the blanks

Checkout crosses [Orders' database], [Inventory's database], and [the payment provider], so no single transaction covers it, and the process can stop between any two steps.
The invariant: every reservation ends in a placed order or is released, and every successful charge belongs to a placed order or is refunded.
Record the checkout in [Orders]' database before anything else, with a key every later step uses: the reservation refers to it, and the provider gets it as the idempotency key. Every step must be safe to repeat.
A provider call that times out is an unknown outcome, not a decline: look the charge up by its key before charging again or giving up.
When the server starts, finish or undo every checkout still pending. A checkout interrupted before the card was charged is undone, not finished: release its stock and tell the customer nothing was charged.
A customer waits at most [10] seconds for an answer; after that the reply is "pending", and recovery keeps working in the background.
Add tests that stop checkout after each step, restart, and assert the invariant over both databases and the provider's records.
Connections to follow nextRelated lessons

Take the shop into your editor. Add a hold that expires after fifteen minutes, and make the audit prove that no expiry ever frees the mug of a pending or placed order.

Back to architecture →