← Concepts & practices
Concept Testing, debugging, and measurement

Diagnosing concurrency bugs

Make the schedule explain the loss.

Two inventory workers decrement the same count. When one worker finishes before the other starts, the result looks right. Under a particular interleaving, both read the same value and the later write erases the earlier update. A sleep may make that failure appear; a controlled schedule and a trace explain why it happens.

TypeScriptGo One counter · two workers · one lost update.

01 / The idea

A passing schedule does not prove a shared update is safe.

The counter starts at 10. Picker A decrements 3; Picker B decrements 4. The expected result is 3. In a sequential schedule, both workers happen to read the latest value.

In the failing schedule, A reads 10 and computes 7. Before A writes, B reads 10 and computes 6. A writes 7; B writes 6. Both operations “ran,” but one update disappeared because the read-modify-write was not one indivisible decision.

Control the interleaving first; then make the shared snapshot visible.

expected310 − 3 − 4
→
observed6picker B overwrote A
The final count shows a lost update. The schedule shows the overlap that caused it.
Read the starting counterTypeScript · fair first answer, unsafe interleaving
inventory.ts · interleaved counter
// A read-modify-write can interleave when two workers share the counter.
export function simulateUnsafe(
	start: number,
	schedule: readonly ScheduleStep[] = unsafeSchedule
): Simulation {
	let available = start;
	const observed: Record<WorkerId, number | null> = { 'picker-a': null, 'picker-b': null };
	const computed: Record<WorkerId, number | null> = { 'picker-a': null, 'picker-b': null };
	const trace: TraceEntry[] = [];
	for (const [index, step] of schedule.entries()) {
		const before = available;
		if (step.action === 'read') {
			observed[step.worker] = available;
			trace.push({
				turn: index + 1,
				...step,
				availableBefore: before,
				availableAfter: available,
				localValue: available,
				note: `${step.worker} reads shared available`
			});
			continue;
		}
		if (step.action === 'compute') {
			const next = (observed[step.worker] ?? available) + jobs[step.worker];
			computed[step.worker] = next;
			trace.push({
				turn: index + 1,
				...step,
				availableBefore: before,
				availableAfter: available,
				localValue: next,
				note: `${step.worker} computes from its snapshot`
			});
			continue;
		}
		available = computed[step.worker] ?? available;
		trace.push({
			turn: index + 1,
			...step,
			availableBefore: before,
			availableAfter: available,
			localValue: available,
			note: `${step.worker} writes its stale snapshot`
		});
	}
	const expected = start + jobs['picker-a'] + jobs['picker-b'];
	return {
		start,
		available,
		expected,
		lostUnits: available - expected,
		trace,
		overlap: firstOverlappingRead(trace) !== null
	};
}

The starting model deliberately separates read, compute, and write. That makes the interleaving explicit without pretending a timer is a synchronization primitive.

02 / See the shape

Separate the schedule, shared state, and worker-local value.

simulateUnsafe follows a schedule of worker actions. Each worker keeps its own observed and computed value while the simulation records the shared counter before and after every turn. simulateSerialized completes one worker’s critical section before the next worker reads.

A serialized critical section: each worker reads, computes, and writes before the next one reads.

TypeScriptReading
inventory.ts
export function simulateSerialized(
	start: number,
	order: readonly WorkerId[] = ['picker-a', 'picker-b']
): Simulation {
	let available = start;
	const trace: TraceEntry[] = [];
	let turn = 0;
	for (const worker of order) {
		const beforeRead = available;
		trace.push({
			turn: ++turn,
			worker,
			action: 'read',
			availableBefore: beforeRead,
			availableAfter: beforeRead,
			localValue: beforeRead,
			note: `${worker} reads while it owns the critical section`
		});
		const next = beforeRead + jobs[worker];
		trace.push({
			turn: ++turn,
			worker,
			action: 'compute',
			availableBefore: beforeRead,
			availableAfter: beforeRead,
			localValue: next,
			note: `${worker} computes before releasing the lock`
		});
		available = next;
		trace.push({
			turn: ++turn,
			worker,
			action: 'write',
			availableBefore: beforeRead,
			availableAfter: available,
			localValue: available,
			note: `${worker} writes before the next worker reads`
		});
	}
	const expected = start + jobs['picker-a'] + jobs['picker-b'];
	return { start, available, expected, lostUnits: available - expected, trace, overlap: false };
}
GoAlongside
inventory.go
func SimulateSerialized(start int, order []WorkerID) Simulation {
	available := start
	trace := []TraceEntry{}
	turn := 0
	for _, worker := range order {
		before := available
		turn++
		trace = append(trace, TraceEntry{Turn: turn, Worker: worker, Action: Read, AvailableBefore: before, AvailableAfter: before, LocalValue: pointer(before), Note: string(worker) + " reads while it owns the critical section"})
		next := before + jobs[worker]
		turn++
		trace = append(trace, TraceEntry{Turn: turn, Worker: worker, Action: Compute, AvailableBefore: before, AvailableAfter: before, LocalValue: pointer(next), Note: string(worker) + " computes before releasing the lock"})
		available = next
		turn++
		trace = append(trace, TraceEntry{Turn: turn, Worker: worker, Action: Write, AvailableBefore: before, AvailableAfter: available, LocalValue: pointer(available), Note: string(worker) + " writes before the next worker reads"})
	}
	expected := start + jobs[PickerA] + jobs[PickerB]
	return Simulation{Start: start, Available: available, Expected: expected, LostUnits: available - expected, Trace: trace, Overlap: false}
}
Reading the TypeScriptControlled actions and trace entries

The model is intentionally deterministic. The worker can be preempted between read, compute, and write; the trace records the local snapshot that later becomes stale.

Reading the GoA model plus a real mutex
inventory.go · mutex version
type SafeInventory struct {
	mu        sync.Mutex
	available int
}

func (inventory *SafeInventory) Adjust(delta int) {
	inventory.mu.Lock()
	defer inventory.mu.Unlock()
	inventory.available += delta
}

func RunSafely(start int, deltas []int) int {
	inventory := &SafeInventory{available: start}
	var workers sync.WaitGroup
	workers.Add(len(deltas))
	for _, delta := range deltas {
		go func() {
			defer workers.Done()
			inventory.Adjust(delta)
		}()
	}
	workers.Wait()
	return inventory.available
}

Go keeps the same explicit schedule and adds a sync.Mutex implementation. The model explains the lost update; go test -race checks that the real synchronized implementation does not introduce a data race.

03 / Follow the schedule

Watch a quiet run become an explainable race.

The frames begin with a schedule that passes by accident, choose the interleaving that loses 3 units, highlight the second worker’s shared read, and serialize the critical section. The count changes; the evidence stays visible.

Concurrency diagnosis

Make the schedule explain the loss.

quiet… available
expected310 − 3 − 4
actual…checking schedule
1picker-a · readlocal 10
2picker-a · computelocal 7
3picker-a · writeshared 7
4picker-b · readlocal 7
5picker-b · computelocal 3
6picker-b · writeshared 3

quiet: actual 3, expected 3; Both updates are present when one worker completes before the next reads.

01/ 04
Run one worker at a time

A quiet schedule passes.

Picker A completes before Picker B reads. The unsafe counter looks correct because the schedule never overlaps the critical section.

Reduced motion: choose a scene to see its completed state.

Read this scene

Picker A completes before Picker B reads. The unsafe counter looks correct because the schedule never overlaps the critical section.

quiet. The race hides under a friendly schedule. Actual 3; expected 3. Both updates are present when one worker completes before the next reads.

Watch restarts when you return. Step through keeps your selected step. Try it starts a fresh schedule.

What controlled diagnosis buys you

A deterministic trace narrows the space of possible causes.

Repeatability
The same action order reaches the same lost count instead of waiting for a lucky timing window.
Shared versus local
The trace distinguishes the current shared value from a worker’s stale computed snapshot.
Small repair
Once the critical section is named, add a lock, atomic operation, channel, or ownership rule at that boundary.
Runtime evidence
The race detector checks real memory access behavior after the controlled model explains the symptom.

Concurrency has more than one failure shape. The next sections distinguish lost updates from ordering, visibility, and liveness problems.

04 / Try a decision

Choose evidence before changing timing.

The inventory count sometimes ends 3 units too high: one picker’s decrement went missing. You can make the race more likely, hide it with retries, or force the read-modify-write overlap. Which move should come first?

The inventory count sometimes ends 3 units too high. What should you do first?

Make the schedule and the read-modify-write state visible before changing synchronization.

05 / Give it a real job

Use the trace to choose synchronization.

First identify the shared state and the critical operation. Then choose the ownership rule that fits: a mutex for a small critical section, an atomic operation for a single supported update, a channel or actor for one owner, or a transactional store for a durable boundary. The trace is diagnostic; it does not decide the architecture for you.

After the fix, replay the controlled schedule, run the real implementation under the race detector or stress harness, and assert the invariant that every worker’s update contributes. Keep the schedule or a close regression scenario where a future reader can find it.

Controlled schedule

Makes the loss repeat

Read, read, write, write, forced in that order, so 10 − 3 − 4 ends at 6 every time instead of once a week.

Synchronization

Gives the count one owner

A mutex around the read-modify-write here. A channel, an atomic, or a transaction where that fits the boundary better.

Race detector

Watches the real run

go test -race over the fixed version, plus a stress run, to catch the interleavings nobody scripted.

Build UIs?Two clicks that read the same count are two workers, and a component that saves totals makes you own the race.

Where it already is in your components

Your stock badge shows a number that several writers change: other shoppers, a warehouse job, an admin in another tab. In the textbook version the component renders the count the server settled on, plus the evidence behind it, so a wrong number points at the server’s schedule, not at the component.

When you have to own it

Now the component does the arithmetic. A “remove 3” button reads available, subtracts, and saves the new total. Two fast clicks, or the same page open in two tabs, read 10 and both save 7. That is the lost update from section 03, and the wild version would show the wrong total with no clue beside it.

Diagnose it the same way. In a test, hold both requests on promises you resolve yourself and release them read, read, write, write. Then fix the boundary: send “remove 3” and let the server apply it, or send the version you read and let the server refuse a stale one.

The UI receives a canonical inventory projection with evidence for an inconsistency.

ReactAlready in your code
textbook.tsx · evidence projection
type InventoryView = {
	available: number;
	expected: number;
	status: 'consistent' | 'lost-update';
	lastEvidence: string;
};

export function InventoryStatus({ inventory }: { inventory: InventoryView }) {
	return (
		<article aria-label="Inventory status">
			<strong>{inventory.available} units available</strong>
			<small>
				{inventory.status === 'consistent'
					? 'Workers agree'
					: `Investigate lost update: expected ${inventory.expected}`}
			</small>
			<p>{inventory.lastEvidence}</p>
		</article>
	);
}

06 / Recognize it elsewhere

Different symptoms point to different concurrency evidence.

Ask what is shared, what is ordered, and what the program promises under interleaving.

Failure shape and first evidence
SymptomLikely shapeFirst evidence
Count too low or highLost update or non-atomic read-modify-writeShared value, worker snapshot, and write order
Older result winsOrdering or stale completionOperation ID, start/finish order, and acceptance rule
One worker sees old dataVisibility or publication gapWrite, synchronization edge, and read observation
Everything waitsDeadlock or blocked dependencyWait-for graph and lock/channel ownership

07 / Already in your toolbox

The tools answer different parts of the question.

A deterministic schedule explains a particular interleaving. A race detector reports unsynchronized memory access in the real program. A stress run explores many schedules. A mutex, atomic, channel, actor, or transaction enforces the rule. Use each tool for the claim it can actually support.

Schedule

Force the overlap

Turn a timing window into a named sequence that can be replayed.

Trace

Show local snapshots

Record shared state, worker identity, and the action that changed it.

Detector

Check the real runtime

Run the synchronized implementation under the race detector and keep the invariant assertion.

08 / The parts to watch

A plausible fix can leave the shared rule unclear.

Sleeping until it failsProbabilistic reproduction

Sleep changes timing, not ownership. Use it as an experiment if needed, but replace it with a controlled schedule for the durable diagnosis.

Locking only the writeThe read is still stale

A lock around assignment does not make read, compute, and write one critical section. Protect the whole operation or use an atomic primitive that expresses it.

Assuming a passing schedule is safeOne order is not all orders

The quiet schedule can pass while the interleaved schedule loses an update. Test allowed outcomes and synchronization guarantees, not one observed order.

Fixing the displayA second truth appears in the UI

Do not clamp or hide a bad count in a component. Fix the owner of the shared state and project the evidence that helps operators understand a remaining failure.

09 / Make the call

Choose a diagnostic depth that matches the failure.

From symptom to synchronization
Start withWhen it helpsThen verify
Controlled scheduleA small interleaving can explain the symptom.Expected versus actual and local snapshots.
Trace or wait graphSeveral workers or dependencies are involved.The ownership or blocking edge.
Race detector / stressThe real runtime can access memory concurrently.Detector-clean execution and the state invariant.
Architecture changeLocks are too broad or ownership is unclear.Single owner, message boundary, or transaction contract.

Make the interleaving repeatable before naming the fix.

Synchronize the whole operation whose correctness depends on a shared snapshot.

Keep this questionAsk it when a race report arrives.

What shared value did each worker read, what local result did it compute, and which synchronization edge was supposed to connect that read to its write?

10 / Take the idea with you

Explain the lost units without saying “race condition.”

“Both workers read 10 before either wrote. A saved 7, then B saved 6 over it, so A’s 3 units disappeared and the count ended 3 too high. We forced that order in a test, put the read-modify-write behind a mutex, and replayed the same order until the count came out at 3.” That tells a reviewer the interleaving, the shared state, and the evidence for the fix. When they want the words, they are lost update, critical section, and controlled schedule.

Before moving on, jot down the four steps that lose the update, why a longer sleep would not prove anything, and one place in your own code where two callers read a value, change it, and write it back. A like counter or a cart total counts.

Connections to follow nextRelated lessons

Take the inventory counter into your editor. Add a third worker that restocks 5 units, then find a schedule that loses the restock instead of a decrement.

Back to Concepts & practices →