← Architecture
Events, queues, and workflows Work for later

Work queues and background jobs

Answer now. Do the work once.

Somewhere in your app a button does slow work: it renders a PDF, resizes an image, or emails a customer, and the request waits for all of it. Moving that work out of the request is easy. Making sure it happens exactly as often as it should, when workers die and providers fail, is the part that needs a design.

TypeScriptGoOne Send button, two designs, a lease you can shorten, two recorded builds.

01 / The prompt

“Send button: render the invoice PDF and email it to the customer. Keep it fast.”

A small invoicing app. Rendering a PDF takes a couple of seconds of work, and the mail provider takes its own time. Done in the request, every Send click waits for both, and a mail outage turns into an error on the customer’s screen.

So the work moves to a background job, and a new set of questions arrives with it. A worker dies halfway: does anyone finish the job? The provider is down for an hour: how many times do you try, and what does the invoice say meanwhile? A job takes longer than expected: does a second worker start the same one? In the lesson’s example, a lease two seconds long for a three-second job ends with the invoice marked failed and the customer emailed 2 times.

The brief never asked the question: once the work leaves the request, what makes sure it happens, and happens once?

02 / Name the shape

The request records a job. A worker leases it, does it, and says so.

A work queue holds jobs: units of work that one worker, any one, should do. The request writes the job in the same transaction as the change that needs it, and answers. A worker takes a job by leasing it for a fixed time, does the work, and acknowledges it. A job whose worker dies comes back when the lease runs out; a job that fails comes back after a backoff; a job that keeps failing stops and is kept for a person.

A job runs at least once and may run twice. Make the lease longer than the work, acknowledge only while you hold it, give the job an attempt limit, and give every outside call the job’s id as its key.

Who owns what:

What each part of sending owns
PartOwnsHands over
The Send requestThe invoice’s status: sendingA job row, in the same commit
The queueEach job’s state, attempts, lease, and last errorOne ready job to one idle worker
A workerIts current job, until its lease runs outAn acknowledgement, or a failure with its error
The mail providerWhether an email went, by keyThe first reply again for a key it has seen

Words to put in a prompt or a review

Job
A unit of work for exactly one worker, recorded before anyone starts it.
Lease
A worker’s claim on a job for a fixed time. Also called a visibility timeout.
Acknowledge
Marking a job done. Only the worker holding the lease may.
Backoff
Waiting longer before each retry, so a failing provider is not hammered.
Dead letter
A job kept aside after its last attempt, with its error, for a person.
Idempotency key
The job’s id sent with each outside call, so running twice has one effect.
A job is not an eventOne worker, versus every reader

An event such as InvoiceSent goes to every reader that cares, each at its own pace; that is Event-driven architecture. A job goes to exactly one worker, and several workers compete for the queue. A queue that many workers read is how a slow task gets done faster; a log that many readers follow is how many tasks hear about one fact.

03 / Send it now, or as a job

Same clicks, same provider. Who waits, and how many emails go?

Each column runs the lesson’s own invoicing code, one click or one second at a time, with two workers. The last chapter compares a sound queue with one whose lease is shorter than the job. Watch, then open Try it and shorten the lease yourself.

Work queues

Send an invoice without waiting for it

Send in the request

The last click waited 3s · mail up

  • inv-1sent1 email

A job for a worker

The last click waited 0s · mail up

  • inv-1sending0 emails
  • w1idle
  • w2idle
  • job-1readyattempt 0
01/ 04
Three invoices

The customer clicks Send on inv-1. Clock: 0s.

Sent in the request, each Send click waits three seconds while the PDF renders and the email goes. As a job, the click is answered at once, and two workers take the invoices in turn.

Reduced motion: choose a scene to see its completed state.

Read this scene

Sent in the request, each Send click waits three seconds while the PDF renders and the email goes. As a job, the click is answered at once, and two workers take the invoices in turn.

Send in the request: inv-1 sent, 1 email(s).

A job for a worker: inv-1 sending, 0 email(s).

Watch restarts the story when you come back. Step through shows where each chapter ends. Try it starts a new app whenever you change a setting or press Reset.

04 / Read the shape

A job written by the request, leased by a worker, finished with a key.

Basic form is the request recording the job. In the wild is the loop every worker runs. At the call site is the worker finishing a job against the mail provider. Notice the check before the acknowledgement.

The request records a job and answers “sending”. The job row is the promise that a worker will render and email the invoice.

TypeScriptReading
app.ts
/**
 * The request records a job and answers. The invoice says "sending" until a
 * worker finishes, and the job row is the promise that one will.
 */
send(request: SendRequest): Reply {
	const invoice = this.create(request);
	this.jobs.push({
		id: `job-${this.jobs.length + 1}`,
		invoice: invoice.id,
		state: 'ready',
		attempts: 0,
		owner: null,
		leaseUntil: null,
		runAfter: this.now,
		error: null
	});
	return { status: 'sending', waited: 0, error: null };
}
GoAlongside
app.go
// Send records a job and answers. The invoice says "sending" until a worker
// finishes, and the job row is the promise that one will.
func (q *QueueApp) Send(r SendRequest) Reply {
	inv := q.create(r)
	q.Jobs = append(q.Jobs, &Job{ID: "job-" + itoa(len(q.Jobs)+1), Invoice: inv.ID, State: "ready", RunAfter: q.Now})
	return Reply{Status: "sending", Waited: 0}
}
Sending in the requestThe first version, for comparison

Render, email, answer. It is correct when nothing fails, and every click pays for all of it.

app.ts
/** The request does the work: render, email, then answer. */
export class InlineApp extends App {
	send(request: SendRequest): Reply {
		const invoice = this.create(request);
		let error = this.render(invoice);
		if (error === null) {
			const result = this.mail.send(null, invoice.id, invoice.email);
			if (result !== 'sent') error = `mail-${result}`;
		}
		invoice.status = error === null ? 'sent' : 'failed';
		return { status: invoice.status, waited: renderSeconds, error };
	}
}
The behavior these examples promiseChecked by 18 shared scenarios
  • In the request, a click waits three seconds, and an outage, a rejected address, or a bad invoice fails the send with no retry.
  • As a job with a six-second lease and a key, a click answers at once; an outage is retried after a backoff; a dead worker’s job is taken over when its lease runs out; a rejected address and a bad invoice fail on the first attempt and are kept.
  • With a two-second lease and no key, every run of a three-second job sends an email, and the job is buried as failed after three attempts.

Every expectation was generated by a separate model written from the contract in the examples’ README, not copied from either implementation, and it is kept beside the examples.

Reading the TypeScriptA clock you move by hand

tick() is one second. The queue lives in the jobs array and each worker in workers; a real queue is a table with the same columns, and tick() is each worker’s polling loop. Object.assign updates several fields of a row at once, the way one UPDATE would.

Reading the GoThe same rows as structs

Jobs and workers are pointers in slices, so a loop can update them in place. A switch with no subject picks the first true case, which reads like the rules: done, buried, or ready again.

Run it yourselfNo dependencies

Save the complete files at the paths in their banners. Then run node --experimental-strip-types run.ts (Node 22.18 or later), or go run . in the Go folder. Both print:

inline · mail down for a while: waited 3s, failed, 0 email(s), - attempt(s) -> consistent
queue · mail down for a while: waited 0s, sent, 1 email(s), 2 attempt(s) -> consistent
short lease · mail down for a while: waited 0s, failed, 1 email(s), 3 attempt(s) -> failed-after-email:inv-1
inline · worker dies mid-job: waited 3s, sent, 1 email(s), - attempt(s) -> consistent
queue · worker dies mid-job: waited 0s, sent, 1 email(s), 2 attempt(s) -> consistent
short lease · worker dies mid-job: waited 0s, failed, 2 email(s), 3 attempt(s) -> emailed-twice:inv-1, failed-after-email:inv-1

05 / Review the agent’s diff

“Some customers got two copies. I complete the job before sending.”

The duplicates are real. Read what the fix trades them for.

The agent’s pull request

“Some customers got two copies of their invoice. It happens when a worker is killed after sending but before completing the job, so the job runs again. I now complete the job before sending. All tests pass.”

// worker.ts
			const job = await queue.lease(workerId, { seconds: 60 });
			const pdf = renderInvoicePdf(invoice);
			(added) // Mark it done first, so a retry can never email the customer twice.
			(added) await queue.complete(job.id);
			await mail.send({ to: invoice.email, attachment: pdf });
			(removed) await queue.complete(job.id);
			
The fix went through code review with two approvals. What do you do?

06 / How it fails

A queue fails quietly: a job that never finishes, or one that finishes twice.

Each row except the last is a shared scenario the tests run.

Failure modes of sending an invoice
What goes wrongIn the requestAs a jobWhat the customer sees
The mail provider is down for a whileThe send fails; nothing retries.Retried after a backoff; sent once the provider is back.“Sending”, then “sent”
The provider never comes backThe send fails.Three attempts, then kept as failed with its error.“Not sent”, with the reason
A worker dies mid-jobNot applicable.The lease runs out, and another worker takes the job.A few seconds’ delay
A rejected address, or an invoice that cannot renderThe send fails.Failed on the first attempt, not retried, and kept.“Not sent”, with the reason
The lease is shorter than the jobNot applicable.A second worker starts every job; without a key, the customer gets every copy.Two emails, then “failed”
Work arrives faster than workers finish itEvery click slows down together.The queue grows; clicks stay fast. Not modeled: add workers, or cap the queue and say so.Invoices arrive later

Retries and duplicates have their own lessons: Retry, backoff, and idempotency and Backpressure and queues.

07 / Is it worth it?

A queue costs a table, a worker process, and a status the page has to show.

The shared four changes, against sending in the request
ChangeIn the requestAs a job
A second entry point: a nightly batch sends overdue remindersThe batch calls the same code and waits for every PDF.The batch records a hundred jobs and ends; workers share them.
A new mail providerChange the call.Change the call in the worker, and check it supports a key.
A new rule: an invoice over a set amount needs a second approverA change before the send.The same change before the job is recorded. No difference.
A team takes over PDF renderingThey change code inside every request path.They own the worker, its deploys, and its scaling.

The costs: a status that says “sending” and a page that has to show it, a worker to deploy and watch, and code that must be safe to run twice. For work that takes 50 ms and cannot fail in a way the user cannot fix, keep it in the request.

Decide what good looks like, and measure it before and after:

  • Send request time at the 95th percentile, which the job should cut to one write.
  • Time from click to email, and the age of the oldest ready job: the queue’s real latency.
  • Emails per sent invoice, exactly one; and dead jobs per day, with their errors.

This lesson did not measure a real app, and gives no numbers.

08 / Ask for it

One brief, two prompts.

Two agents running Claude Sonnet each got the brief from section 01 and a render.ts that spins the CPU for as long as it is told. One prompt added a Jobs block: a job recorded with the status change, a lease longer than the work, at most five attempts with backoff, and the job’s id as the email’s key. A script ran both builds with real worker processes against its own mail provider, and killed workers and servers along the way.

What the checker found, run 2026-09-23
QuestionPlain promptJobs prompt
Ten invoices sent at onceReplies in 3 ms or less; 10 sent, 1 email eachReplies in 4 ms or less; 10 sent, 1 email each
Every worker killed mid-job, a new one startedStill sending after 90 ssent after 33 s, 1 email
The mail provider down for 20 secondsDuring: 2 sending. After: 2 sent. Attempts: 40 and 1During: 2 failed. After: 2 failed. Attempts: 5 and 5
An address the provider rejectsfailed after 1 attempt; the other sentfailed after 1 attempt; the other sent
The same invoice sent twice1 email1 email
The provider sends, then drops the connection1 email each1 email each
Server and worker killed with sends outstanding1 sending, 2 sent after 60 s3 sent after 30 s
Its own tests27 of 27 pass18 of 18 pass

Both agents built a queue in the database, with a claim that expires and a key on every email. The brief asked for a 200 ms answer and for worker processes, and that was enough to get the shape. Neither build sent a second email in any question.

The difference was time, and neither prompt gave any. The plain build’s claim lasts two minutes, so an invoice held by a dead worker sat in “sending” past the checker’s 90 seconds. Its retries have no pause and no limit: during a 20-second outage it tried the first invoice 40 times and the second once, because the failing one was claimed again at every poll.

worker.ts and worker-lib.ts · plain prompt
const LEASE_MS = Number(process.env.WORKER_LEASE_MS ?? 120_000);
…
  if (result.permanent) {
    deps.store.markFailed(invoice.id);
    return 'failed-rejected';
  }
  deps.store.release(invoice.id);
  return 'retry';

The jobs build did exactly what its block said: five attempts, with backoff. Its backoff starts at half a second and doubles, so the five attempts are used up in about eight seconds, and a 20-second outage left both invoices failed for good, with nothing sent.

worker.ts and db.ts · jobs prompt
    backoffBaseMs: Number(process.env.JOB_BACKOFF_BASE_MS ?? 500),
…
export function backoffForAttempt(attempt: number, baseMs: number): number {
  // attempt 1 -> baseMs, attempt 2 -> 2x, attempt 3 -> 4x, ...
  return baseMs * 2 ** (attempt - 1);
}
…
    // Transient error: retry with backoff, unless attempts are exhausted.
    if (job.attempts >= job.max_attempts) {
      markJobFailed(db, job.id, result.error);
    } else {
      markJobRetry(db, job.id, result.error, backoffForAttempt(job.attempts, config.backoffBaseMs));
    }

The missing lines are numbers with a reason: the lease is a little longer than the slowest normal job, and retries span the longest outage you want to ride out, say an hour, before a job is kept as failed. An attempt count without a time scale is a guess.

How the runs were made and checkedTwo builds, recorded as written
  • Both agents started from the same render.ts and were launched at the same time; neither was told about the other, the lesson, or the checker. Neither changed render.ts.
  • Both builds are kept byte for byte with checksums. For every question the checker restores a build into a fresh folder with its own database, starts the server and worker processes, and runs its own mail provider.
  • The checker ran twice; the first covered the jobs build alone and gave the same answers. The plain build’s two-minute claim was not waited out; that finding is read from its code.
  • The jobs agent created and deleted one file in /tmp. Neither stopped a process by name or pattern.
  • One run of each prompt is a sample, not a measurement of the model.

09 / Hold it there

The rules of a queue are easy to write and easy to lose. Three checks keep them.

  1. The queue’s own door: the lease, and at-least-once

    Every hosted queue has the same two facts in its documentation; know them for yours. Amazon SQS calls the lease a visibility timeout: “If you don’t delete it before the timeout expires, the message becomes visible again in the queue and can be retrieved by another consumer. The default visibility timeout for a queue is 30 seconds.” And it does not promise once: “because of the at-least-once delivery model, Amazon SQS doesn’t guarantee that a message won’t be delivered more than once within the visibility timeout period.” (Amazon SQS visibility timeout) A default lease of 30 seconds and a job that takes 45 is the short-lease chapter, in production.

  2. Tests that kill a worker and shorten the lease

    The shared scenarios kill a worker mid-job, take the provider down past the last attempt, and run a lease shorter than the job; each ends with a count of emails per invoice. A rule that every call leaving a worker carries a key can be enforced the way Enforcement layer enforces import rules.

  3. Watch the queue itself

    Alert on the age of the oldest ready job, on jobs kept as failed, and on emails sent per invoice. A queue that looks empty because its worker is dead is the one to catch first.

Build UIs?Every “Processing…” you have shown was a job; the question is whether it was honest.

Where it already is in your components

An Export button that says “Preparing your file…” and later offers a download is a job’s status on screen. So is an upload that says “Processing” while a video is transcoded. The component starts the work, then reports what the server says about it.

When you have to own it

When the server’s status can be “sending”, “retrying”, or “failed”, the page must show which, and must not turn a job it merely started into “Sent!”. Poll or subscribe while the job is open, and show the error when it is kept as failed.

An export button that starts a job and polls it until it is done or failed.

ReactAlready in your code
ExportButton.tsx
import { useState } from 'react';

type Job = {
	id: string;
	state: 'queued' | 'running' | 'done' | 'failed';
	url?: string;
	error?: string;
};

// An export that takes minutes is a job, not a request. The button starts it,
// then shows what the server says about it, and never pretends it is done.
export function ExportButton() {
	const [job, setJob] = useState<Job | null>(null);

	async function start() {
		const started = (await (await fetch('/api/exports', { method: 'POST' })).json()) as Job;
		setJob(started);
		let current = started;
		while (current.state === 'queued' || current.state === 'running') {
			await new Promise((resolve) => setTimeout(resolve, 2000));
			current = (await (await fetch(`/api/jobs/${started.id}`)).json()) as Job;
			setJob(current);
		}
	}

	if (!job)
		return (
			<button type="button" onClick={start}>
				Export invoices
			</button>
		);
	if (job.state === 'done') return <a href={job.url}>Download the export</a>;
	if (job.state === 'failed') return <p role="alert">The export failed: {job.error}</p>;
	return <p role="status">{job.state === 'queued' ? 'Waiting to start…' : 'Exporting…'}</p>;
}

10 / Make the call

Move slow or fallible work to a job, and design for it running twice.

Keep work in the request when it is quick and its failure is something the user can fix on the spot. Make it a job when it is slow, depends on a service that fails, or must survive a restart; then give the job a lease longer than the work, an attempt limit, a place to rest when it fails for good, and a key for every outside call. Reopen the decision when the job needs several steps that must all happen or be undone; that is a saga.

Take it with you

Explain it without saying “queue”: “Clicking Send writes a ticket. A worker picks the ticket up, holds it for a few minutes, and does the work. If the worker disappears, the ticket goes back on the pile, and the work is written so that doing it twice sends one email.” Then find the slowest thing one of your requests does after its main write, and ask what would happen if the process died right there.

Paste into your next prompt, and fill in the blanks

[Sending an invoice] is a job, recorded in [the app's database] in the same transaction that sets [the invoice] to [sending]. The request answers after that commit.
Workers lease a job for [a little longer than the slowest normal job]; while the lease holds no other worker takes it, and when a worker dies the job is taken again once the lease runs out. A worker acknowledges a job only while it still holds the lease.
A failed job is retried with backoff that spans [the longest outage to ride out, such as an hour], at most [N] attempts. After the last, [the invoice] is [failed] and the job is kept with its error for a person to look at. [A rejected address] is not retried.
A job may run more than once: every outside call it makes carries the job's id as its idempotency key.
Expose the queue's depth, the oldest ready job's age, and the jobs kept as failed.
Connections to follow nextRelated lessons

Take the queue into your editor. Add a heartbeat that extends a worker’s lease while it is still rendering, and make the short-lease chapter end with one email.

Back to architecture →