01 / The idea
A passing schedule does not prove a shared update is safe.
The counter starts at 10. Picker A decrements 3; Picker B decrements 4. The expected result is 3. In a sequential schedule, both workers happen to read the latest value.
In the failing schedule, A reads 10 and computes 7. Before A writes, B reads 10 and computes 6. A writes 7; B writes 6. Both operations “ran,” but one update disappeared because the read-modify-write was not one indivisible decision.
Control the interleaving first; then make the shared snapshot visible.
Read the starting counterTypeScript · fair first answer, unsafe interleaving
// A read-modify-write can interleave when two workers share the counter.
export function simulateUnsafe(
start: number,
schedule: readonly ScheduleStep[] = unsafeSchedule
): Simulation {
let available = start;
const observed: Record<WorkerId, number | null> = { 'picker-a': null, 'picker-b': null };
const computed: Record<WorkerId, number | null> = { 'picker-a': null, 'picker-b': null };
const trace: TraceEntry[] = [];
for (const [index, step] of schedule.entries()) {
const before = available;
if (step.action === 'read') {
observed[step.worker] = available;
trace.push({
turn: index + 1,
...step,
availableBefore: before,
availableAfter: available,
localValue: available,
note: `${step.worker} reads shared available`
});
continue;
}
if (step.action === 'compute') {
const next = (observed[step.worker] ?? available) + jobs[step.worker];
computed[step.worker] = next;
trace.push({
turn: index + 1,
...step,
availableBefore: before,
availableAfter: available,
localValue: next,
note: `${step.worker} computes from its snapshot`
});
continue;
}
available = computed[step.worker] ?? available;
trace.push({
turn: index + 1,
...step,
availableBefore: before,
availableAfter: available,
localValue: available,
note: `${step.worker} writes its stale snapshot`
});
}
const expected = start + jobs['picker-a'] + jobs['picker-b'];
return {
start,
available,
expected,
lostUnits: available - expected,
trace,
overlap: firstOverlappingRead(trace) !== null
};
} The starting model deliberately separates read, compute, and write. That makes the interleaving explicit without pretending a timer is a synchronization primitive.
02 / See the shape
Separate the schedule, shared state, and worker-local value.
simulateUnsafe follows a schedule of worker actions. Each worker keeps its own
observed and computed value while the simulation records the shared counter before and after
every turn. simulateSerialized completes one worker’s critical section before the
next worker reads.
A serialized critical section: each worker reads, computes, and writes before the next one reads.
export function simulateSerialized(
start: number,
order: readonly WorkerId[] = ['picker-a', 'picker-b']
): Simulation {
let available = start;
const trace: TraceEntry[] = [];
let turn = 0;
for (const worker of order) {
const beforeRead = available;
trace.push({
turn: ++turn,
worker,
action: 'read',
availableBefore: beforeRead,
availableAfter: beforeRead,
localValue: beforeRead,
note: `${worker} reads while it owns the critical section`
});
const next = beforeRead + jobs[worker];
trace.push({
turn: ++turn,
worker,
action: 'compute',
availableBefore: beforeRead,
availableAfter: beforeRead,
localValue: next,
note: `${worker} computes before releasing the lock`
});
available = next;
trace.push({
turn: ++turn,
worker,
action: 'write',
availableBefore: beforeRead,
availableAfter: available,
localValue: available,
note: `${worker} writes before the next worker reads`
});
}
const expected = start + jobs['picker-a'] + jobs['picker-b'];
return { start, available, expected, lostUnits: available - expected, trace, overlap: false };
} func SimulateSerialized(start int, order []WorkerID) Simulation {
available := start
trace := []TraceEntry{}
turn := 0
for _, worker := range order {
before := available
turn++
trace = append(trace, TraceEntry{Turn: turn, Worker: worker, Action: Read, AvailableBefore: before, AvailableAfter: before, LocalValue: pointer(before), Note: string(worker) + " reads while it owns the critical section"})
next := before + jobs[worker]
turn++
trace = append(trace, TraceEntry{Turn: turn, Worker: worker, Action: Compute, AvailableBefore: before, AvailableAfter: before, LocalValue: pointer(next), Note: string(worker) + " computes before releasing the lock"})
available = next
turn++
trace = append(trace, TraceEntry{Turn: turn, Worker: worker, Action: Write, AvailableBefore: before, AvailableAfter: available, LocalValue: pointer(available), Note: string(worker) + " writes before the next worker reads"})
}
expected := start + jobs[PickerA] + jobs[PickerB]
return Simulation{Start: start, Available: available, Expected: expected, LostUnits: available - expected, Trace: trace, Overlap: false}
} Reading the TypeScriptControlled actions and trace entries
The model is intentionally deterministic. The worker can be preempted between read, compute, and write; the trace records the local snapshot
that later becomes stale.
Reading the GoA model plus a real mutex
type SafeInventory struct {
mu sync.Mutex
available int
}
func (inventory *SafeInventory) Adjust(delta int) {
inventory.mu.Lock()
defer inventory.mu.Unlock()
inventory.available += delta
}
func RunSafely(start int, deltas []int) int {
inventory := &SafeInventory{available: start}
var workers sync.WaitGroup
workers.Add(len(deltas))
for _, delta := range deltas {
go func() {
defer workers.Done()
inventory.Adjust(delta)
}()
}
workers.Wait()
return inventory.available
} Go keeps the same explicit schedule and adds a sync.Mutex implementation. The model explains the lost update; go test -race checks that
the real synchronized implementation does not introduce a data race.
03 / Follow the schedule
Watch a quiet run become an explainable race.
The frames begin with a schedule that passes by accident, choose the interleaving that loses 3 units, highlight the second worker’s shared read, and serialize the critical section. The count changes; the evidence stays visible.
Make the schedule explain the loss.
picker-a · readlocal 10picker-a · computelocal 7picker-a · writeshared 7picker-b · readlocal 7picker-b · computelocal 3picker-b · writeshared 3quiet: actual 3, expected 3; Both updates are present when one worker completes before the next reads.
A quiet schedule passes.
Picker A completes before Picker B reads. The unsafe counter looks correct because the schedule never overlaps the critical section.
Reduced motion: choose a scene to see its completed state.
Read this scene
Picker A completes before Picker B reads. The unsafe counter looks correct because the schedule never overlaps the critical section.
quiet. The race hides under a friendly schedule. Actual 3; expected 3. Both updates are present when one worker completes before the next reads.
Watch restarts when you return. Step through keeps your selected step. Try it starts a fresh schedule.
What controlled diagnosis buys you
A deterministic trace narrows the space of possible causes.
- Repeatability
- The same action order reaches the same lost count instead of waiting for a lucky timing window.
- Shared versus local
- The trace distinguishes the current shared value from a worker’s stale computed snapshot.
- Small repair
- Once the critical section is named, add a lock, atomic operation, channel, or ownership rule at that boundary.
- Runtime evidence
- The race detector checks real memory access behavior after the controlled model explains the symptom.
Concurrency has more than one failure shape. The next sections distinguish lost updates from ordering, visibility, and liveness problems.
04 / Try a decision
Choose evidence before changing timing.
The inventory count sometimes ends 3 units too high: one picker’s decrement went missing. You can make the race more likely, hide it with retries, or force the read-modify-write overlap. Which move should come first?
05 / Give it a real job
Use the trace to choose synchronization.
First identify the shared state and the critical operation. Then choose the ownership rule that fits: a mutex for a small critical section, an atomic operation for a single supported update, a channel or actor for one owner, or a transactional store for a durable boundary. The trace is diagnostic; it does not decide the architecture for you.
After the fix, replay the controlled schedule, run the real implementation under the race detector or stress harness, and assert the invariant that every worker’s update contributes. Keep the schedule or a close regression scenario where a future reader can find it.
Makes the loss repeat
Read, read, write, write, forced in that order, so 10 − 3 − 4 ends at 6 every time instead of once a week.
Gives the count one owner
A mutex around the read-modify-write here. A channel, an atomic, or a transaction where that fits the boundary better.
Watches the real run
go test -race over the fixed version, plus a stress run, to catch the interleavings
nobody scripted.
Build UIs?Two clicks that read the same count are two workers, and a component that saves totals makes you own the race.
Where it already is in your components
Your stock badge shows a number that several writers change: other shoppers, a warehouse job, an admin in another tab. In the textbook version the component renders the count the server settled on, plus the evidence behind it, so a wrong number points at the server’s schedule, not at the component.
When you have to own it
Now the component does the arithmetic. A “remove 3” button reads available,
subtracts, and saves the new total. Two fast clicks, or the same page open in two tabs,
read 10 and both save 7. That is the lost update from section 03, and the wild version
would show the wrong total with no clue beside it.
Diagnose it the same way. In a test, hold both requests on promises you resolve yourself and release them read, read, write, write. Then fix the boundary: send “remove 3” and let the server apply it, or send the version you read and let the server refuse a stale one.
The UI receives a canonical inventory projection with evidence for an inconsistency.
type InventoryView = {
available: number;
expected: number;
status: 'consistent' | 'lost-update';
lastEvidence: string;
};
export function InventoryStatus({ inventory }: { inventory: InventoryView }) {
return (
<article aria-label="Inventory status">
<strong>{inventory.available} units available</strong>
<small>
{inventory.status === 'consistent'
? 'Workers agree'
: `Investigate lost update: expected ${inventory.expected}`}
</small>
<p>{inventory.lastEvidence}</p>
</article>
);
}
06 / Recognize it elsewhere
Different symptoms point to different concurrency evidence.
Ask what is shared, what is ordered, and what the program promises under interleaving.
| Symptom | Likely shape | First evidence |
|---|---|---|
| Count too low or high | Lost update or non-atomic read-modify-write | Shared value, worker snapshot, and write order |
| Older result wins | Ordering or stale completion | Operation ID, start/finish order, and acceptance rule |
| One worker sees old data | Visibility or publication gap | Write, synchronization edge, and read observation |
| Everything waits | Deadlock or blocked dependency | Wait-for graph and lock/channel ownership |
07 / Already in your toolbox
The tools answer different parts of the question.
A deterministic schedule explains a particular interleaving. A race detector reports unsynchronized memory access in the real program. A stress run explores many schedules. A mutex, atomic, channel, actor, or transaction enforces the rule. Use each tool for the claim it can actually support.
Force the overlap
Turn a timing window into a named sequence that can be replayed.
Show local snapshots
Record shared state, worker identity, and the action that changed it.
Check the real runtime
Run the synchronized implementation under the race detector and keep the invariant assertion.
08 / The parts to watch
A plausible fix can leave the shared rule unclear.
Sleeping until it failsProbabilistic reproduction
Sleep changes timing, not ownership. Use it as an experiment if needed, but replace it with a controlled schedule for the durable diagnosis.
Locking only the writeThe read is still stale
A lock around assignment does not make read, compute, and write one critical section. Protect the whole operation or use an atomic primitive that expresses it.
Assuming a passing schedule is safeOne order is not all orders
The quiet schedule can pass while the interleaved schedule loses an update. Test allowed outcomes and synchronization guarantees, not one observed order.
Fixing the displayA second truth appears in the UI
Do not clamp or hide a bad count in a component. Fix the owner of the shared state and project the evidence that helps operators understand a remaining failure.
09 / Make the call
Choose a diagnostic depth that matches the failure.
| Start with | When it helps | Then verify |
|---|---|---|
| Controlled schedule | A small interleaving can explain the symptom. | Expected versus actual and local snapshots. |
| Trace or wait graph | Several workers or dependencies are involved. | The ownership or blocking edge. |
| Race detector / stress | The real runtime can access memory concurrently. | Detector-clean execution and the state invariant. |
| Architecture change | Locks are too broad or ownership is unclear. | Single owner, message boundary, or transaction contract. |
Make the interleaving repeatable before naming the fix.
Synchronize the whole operation whose correctness depends on a shared snapshot.
Keep this questionAsk it when a race report arrives.
What shared value did each worker read, what local result did it compute, and which synchronization edge was supposed to connect that read to its write?
10 / Take the idea with you
Explain the lost units without saying “race condition.”
“Both workers read 10 before either wrote. A saved 7, then B saved 6 over it, so A’s 3 units disappeared and the count ended 3 too high. We forced that order in a test, put the read-modify-write behind a mutex, and replayed the same order until the count came out at 3.” That tells a reviewer the interleaving, the shared state, and the evidence for the fix. When they want the words, they are lost update, critical section, and controlled schedule.
Before moving on, jot down the four steps that lose the update, why a longer sleep would not prove anything, and one place in your own code where two callers read a value, change it, and write it back. A like counter or a cart total counts.
Connections to follow nextRelated lessons
- Event loop vs. scheduler explains why runtime scheduling does not provide the synchronization your state needs.
- Race conditions in UI follows the same kind of bug through overlapping requests in a component.
- Optimistic concurrency refuses a stale write with a version check instead of a lock.
- Debugging state and control flow provides the same reproduce, trace, and replay discipline for sequential workflows.
- Invariants and example tests names the count rule the fix has to keep.