← Architecture
Failure and evidence Before and after

Baseline before you change

Measure it, then change it.

Every performance change you have merged came with a feeling: the page snaps now, the job finishes sooner. Sometimes the feeling was right. Let’s take one refactor that really is faster, and find out what else it changed that nobody would have felt.

The skill to keep: Before a change, fix the workload, record exactly what it outputs, time it enough times to see the spread, and write down what result you would accept. After the change, answer that same question, behavior first.

TypeScriptGoOne report, three versions, two recorded refactors.

01 / The prompt

“The weekly report is slow. Make it faster.”

A food bank prints a report every week: each household, how many times it came, and how many pounds of food it took home, heaviest first. With 1,200 households it has started to drag. You ask an agent to speed it up, and it comes back with a clean one-pass rewrite and a message: much faster now. You run it. It is.

Nobody timed the old version before the change, so “much faster” is a feeling with a direction. More importantly, nobody saved what the old version printed. The report is read by volunteers who decide who gets a delivery, so a household that drops off the list, or a total that moves by a tenth of a pound, is not a detail.

The question the prompt never answered: faster than what, on which data, and printing the same thing? Without a baseline taken before the change, nothing after it can answer.

02 / Name the move

Four things, written down before you touch the code.

A baseline is a record of how the system behaves and performs before a change, taken so that the change can be judged against it. It has four parts, and the order matters: the output comes before the timing, because a faster report that prints something else is a different report.

Fix the workload. Snapshot the output. Measure more than once. Decide what you would accept. Then change it, and answer the same question.

The four parts of a baseline, for the food bank’s report
PartFor this reportWithout it
A fixed workloadWeek 38: 1,200 households, 6,000 pickups, and a few edge cases.Before and after ran on different data, and the difference means nothing.
An output snapshotEvery line printed, and its fingerprint, 9a75b00e.A behavior change looks like a speedup.
Repeated samplesFifteen timed runs after three warmups, every sample kept.One lucky or cold run decides.
A rule, decided firstSame fingerprint; then every run after beats every run before.The rule bends to fit whatever the numbers say.

Words to put in a prompt or a review

Baseline
The before picture: behavior and performance, on a fixed workload.
Workload
The exact input both measurements run on.
Snapshot
What the code outputs today, saved so tomorrow’s output can be compared.
Warmup
Runs thrown away so caches and the JIT settle before timing starts.
Spread
How much repeated runs differ. A change smaller than the spread is not a result.
Acceptance rule
What would count as better, worse, or the same, written before the change.
Why count steps as well as timeA number the machine cannot move

Wall-clock time depends on the machine, what else it is doing, and the runtime’s mood. A count of the work, here the inner-loop steps, does not: the original takes 7,200,000 on week 38 on every machine, the refactor 6,000. Counts explain why a change is faster; timings say whether it is, on a real machine. Keep both.

03 / Follow one refactor

Timing says yes. The snapshot says look again.

First the refactor judged on timing alone. Then the same refactor with the output snapshot taken first. Then the fix, judged by the same rule. Open Try it to take a real baseline in your browser.

Failure and evidence

Was the refactor worth it?

  1. 1Fix the workload
  2. 2Snapshot the output
  3. 3Measure, more than once
  4. 4Apply the rule

Week 38 · 1,200 households, 6,000 pickups

Steps before
7,200,000
Steps · the agent’s refactor
6,000
Fingerprint before
9a75b00e
Fingerprint · the agent’s refactor
not taken

Timed runs

Not measured yet.

Verdict: —

01/ 03
It felt faster

“Made the report much faster: one pass instead of a loop inside a loop.”

The agent is right about the work: 7,200,000 inner-loop steps before, 6,000 after, on week 38.

Reduced motion: choose a scene to see its completed state.

Read this scene

The agent is right about the work: 7,200,000 inner-loop steps before, 6,000 after, on week 38.

The agent’s refactor. Steps: 7200000 before, 6000 after. Fingerprint before 9a75b00e, after not taken. Verdict: not decided.

Watch restarts the story when you come back. Step through keeps your step. Try it takes a real baseline in your browser.

04 / Read it in code

A rule, a measurement, and a record.

Basic form is the verdict, written before any number exists. In the wild is the measurement, with warmup and a check that every run printed the same thing. At the call site is the record you keep, with the rule stored next to the numbers so the after-measurement answers the same question.

The rule, decided before the change: a different output fingerprint means the change is judged on behavior, not speed. Then “faster” only if every run after beat every run before; overlap is inconclusive.

TypeScriptReading
baseline.ts
export type Measure = { fingerprint: string; lines: number; samples: number[] };
export type Verdict =
	| { kind: 'behavior-changed'; missing: number; changed: number }
	| { kind: 'faster' | 'slower' | 'inconclusive'; medianRatio: number };

export const median = (samples: readonly number[]) => {
	const s = [...samples].sort((a, b) => a - b);
	const mid = Math.floor(s.length / 2);
	return s.length % 2 ? s[mid] : (s[mid - 1] + s[mid]) / 2;
};

/**
 * The rule, decided before the change: the output must be the same, and "faster" means the
 * slowest run after the change beat the fastest run before it. Overlap is not a result.
 */
export function verdict(
	before: Measure,
	after: Measure,
	diff: { missing: number; changed: number }
): Verdict {
	if (before.fingerprint !== after.fingerprint) return { kind: 'behavior-changed', ...diff };
	const medianRatio = Math.round((median(after.samples) / median(before.samples)) * 100) / 100;
	if (Math.max(...after.samples) < Math.min(...before.samples))
		return { kind: 'faster', medianRatio };
	if (Math.min(...after.samples) > Math.max(...before.samples))
		return { kind: 'slower', medianRatio };
	return { kind: 'inconclusive', medianRatio };
}
GoAlongside
main.go
type Measure struct {
	Fingerprint string
	Lines       int
	Samples     []float64
}

// Verdict is "behavior-changed", "faster", "slower", or "inconclusive".
type Verdict struct {
	Kind        string  `json:"kind"`
	Missing     int     `json:"missing,omitempty"`
	Changed     int     `json:"changed,omitempty"`
	MedianRatio float64 `json:"medianRatio,omitempty"`
}

func Median(samples []float64) float64 {
	s := append([]float64(nil), samples...)
	sort.Float64s(s)
	mid := len(s) / 2
	if len(s)%2 == 1 {
		return s[mid]
	}
	return (s[mid-1] + s[mid]) / 2
}

func minMax(s []float64) (float64, float64) {
	lo, hi := s[0], s[0]
	for _, v := range s {
		lo, hi = math.Min(lo, v), math.Max(hi, v)
	}
	return lo, hi
}

// Decide applies the rule decided before the change: the output must be the same, and
// "faster" means the slowest run after the change beat the fastest run before it.
func Decide(before, after Measure, missing, changed int) Verdict {
	if before.Fingerprint != after.Fingerprint {
		return Verdict{Kind: "behavior-changed", Missing: missing, Changed: changed}
	}
	ratio := math.Round(Median(after.Samples)/Median(before.Samples)*100) / 100
	beforeMin, beforeMax := minMax(before.Samples)
	afterMin, afterMax := minMax(after.Samples)
	switch {
	case afterMax < beforeMin:
		return Verdict{Kind: "faster", MedianRatio: ratio}
	case afterMin > beforeMax:
		return Verdict{Kind: "slower", MedianRatio: ratio}
	}
	return Verdict{Kind: "inconclusive", MedianRatio: ratio}
}
The behavior these examples promiseChecked by 15 shared cases in TypeScript and Go
  • The report lists every household with its visits and pounds, summed in hundredths and rounded once to tenths, heaviest first, ties by name. The agent’s refactor rounds each pickup first and lists only households that came; the fixed refactor keeps the original’s behavior in one pass.
  • The fingerprint is FNV-1a over the report’s text; the diff counts households missing and printed differently.
  • The verdict: a different fingerprint is “behavior changed”. Otherwise “faster” when the slowest run after is below the fastest run before, “slower” in the mirror case, and “inconclusive” when they overlap.

Every expectation in the shared cases was produced by a separate model written from these rules, kept beside the examples in examples/model/, not copied from either implementation: three workloads in three versions each, and six verdicts on given samples, including overlap and a single run each.

Reading the TypeScriptAn injected clock and integer pounds

measure takes its clock as an option, so the tests can count calls with a fake one and the lab can use performance.now(). Pounds are integer hundredths, and rounding is Math.floor((x + 5) / 10), so the three languages that check this code agree to the last digit.

Reading the GoA function type and time.Now

Report is a function type, so the three versions and the flaky one in the tests are interchangeable. MeasureReport takes now func() time.Time for the same reason measure takes a clock. Go’s own benchmark harness, go test -bench, repeats runs for you; the lesson writes the loop out so the rule is visible.

Run it yourselfNo dependencies

Copy the complete TypeScript file and run node --experimental-strip-types baseline.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .; add measure to take baseline records too. Both print:

go.mod
module heyrian.dev/lessons/baseline-before-change

go 1.23
before: 1200 lines, fingerprint 9a75b00e, 7200000 steps
after: 1176 lines, fingerprint f64824a6, 6000 steps; 24 missing, 535 changed
fixed: 1200 lines, fingerprint 9a75b00e, 7200 steps; 0 missing, 0 changed

05 / Review the agent’s diff

“It is much faster now.”

The diagnosis is right: a loop over every pickup inside a loop over every household is the slow part. Read what the rewrite prints before you decide how to check it.

The agent’s pull request

“Rewrote the weekly report to group pickups in one pass instead of scanning every pickup for every household. It is much faster now.”

// report.ts
			(removed)const rows = week.households.map((household) => {
			(removed)  for (const pickup of week.pickups) { /* sum this household */ }
			(added)const totals = new Map<string, { visits: number; tenths: number }>();
			(added)for (const pickup of week.pickups) {
			(added)  const t = totals.get(pickup.household) ?? { visits: 0, tenths: 0 };
			(added)  t.visits++;
			(added)  t.tenths += Math.round(pickup.hundredths / 10);
			(added)  totals.set(pickup.household, t);
			(added)}
			(added)const rows = [...totals].map(([id, t]) => ({ name: names.get(id)!, ...t }));
			
You are reviewing this change. What do you do?

06 / How it fails

A baseline fails by measuring the wrong thing, or not enough of it.

Each way a before-and-after comparison can mislead, what it would have told you, and what the procedure does instead. The numbers come from the shared cases and the captured run.

How a before-and-after comparison misleads
What goes wrongWhat it would have saidWhat the procedure doesFrom the cases
Timing without an output check“Much faster.”Fingerprint first; timing only for the same output.after: 24 missing, 535 changed
One run eachWhatever that run said.Fifteen runs, all kept.single run: “faster”
The runs overlap“A bit faster” from the medians.Inconclusive, and says so.overlap: median ratio 0.91, “inconclusive”
One slow outlier“No faster.”Inconclusive under this strict rule; look at the samples.outlier: “inconclusive”
Only the data you tested with“Same output” on week 38.Add edge cases: nobody came, rounding, ties, an unknown id.small workload: 1 missing, 1 changed
A cold first runA slower “before” than is fair.Warmup runs, thrown away.tested with a fake clock

The strict rule, no overlap at all, is a choice, and it is the one this lesson makes because the refactor it cares about changes the work by a factor of a thousand. For a change you expect to move things by five percent, decide a statistical test in advance instead; the point is that the rule exists before the numbers do. Deciding what a user would count as success is Defining success; finding where the time goes is CPU and memory profiling.

07 / Is it worth it?

You pay a few minutes and a file. Here is what they buy.

Skipping the baseline is faster, and most of the time the change is fine. Hold both habits up against the changes this report will get.

The same four changes, with and without a baseline
ChangeNo baselineBaseline on record
A second client: a CSV export for the cityA second output nobody snapshotted.Snapshot it too; the same record covers both.
Replace the JSON reader with a streaming one“Seems fine.”Same fingerprint and the timings, or it does not merge.
Change a rule: round to whole poundsIndistinguishable from a bug.The snapshot changes on purpose, and the new one becomes the baseline.
A new volunteer maintains itNo idea what “normal” was.A file that says what it printed, how fast, and on what.

The technique is itself a measurement plan, so this section’s plan is for the habit: over the next five performance changes you merge, count how many came with a baseline, and how many of those found a behavior change or an unconvincing speedup. Keep the habit where either number is more than zero.

The captured timings in this lesson come from one machine on one afternoon (Apple M3 Max, median 33.473 ms before, 0.456 ms after the fix). They show a distribution; they are not a benchmark.

08 / Ask for it

Two refactors, both correct. One of them can prove it.

We gave two agents, both running Claude Sonnet, the same slow report and week 38’s data, and asked them to make it faster. One prompt stopped there. The other added a paragraph: take a baseline first, record what the report prints, time it with repeated runs, decide what you would accept, measure the same way after, and write it all to BASELINE.md. Then a script ran each refactor beside the original on weeks the agents never saw.

What each refactor printed beside the original, run 2026-09-23
WorkloadPlain promptArchitecture prompt
Week 38 (the agents’ own data)Same outputSame output
Week 39 (another week)Same outputSame output
Edge cases (a household that never came, rounding, a tie, an unknown household)Same outputSame output
An empty weekSame outputSame output
A week four times larger, median of eight runs353 → 75 ms, no overlap363 → 75 ms, no overlap
A written baselineNoneBASELINE.md

The result was not the one this lesson’s own example predicts. Both agents found the loop inside a loop, replaced it with one pass, and kept the report exactly as it was, on every workload including the edge cases. Both are about as fast as each other.

The difference is what each left behind. The plain agent did half a baseline without being asked: it saved the output before and after, diffed them, then deleted both files and summed up the speed as “the difference isn’t visible yet” at this size. The architecture agent left a file a reviewer can check: the workload, the output’s hash, twenty-five timed runs, a rule it wrote before the change, and the result against that rule. It even measured Node’s startup, about 37 of its 63 milliseconds, so the numbers would not overstate the report’s own time.

report.ts · plain prompt
const totals = new Map<string, { hundredths: number; visits: number }>();
for (const pickup of week.pickups) {
	const totalsForHousehold = totals.get(pickup.household);
	if (totalsForHousehold) {
		totalsForHousehold.hundredths += pickup.hundredths;
		totalsForHousehold.visits++;
	} else {
		totals.set(pickup.household, { hundredths: pickup.hundredths, visits: 1 });
	}
}

const rows = week.households.map((household) => {
	const totalsForHousehold = totals.get(household.id);
	const hundredths = totalsForHousehold ? totalsForHousehold.hundredths : 0;
	const visits = totalsForHousehold ? totalsForHousehold.visits : 0;
	return { name: household.name, visits, tenths: Math.floor((hundredths + 5) / 10) };
});
BASELINE.md · architecture prompt
## The rule, fixed before touching the code

Decided from the baseline's spread alone (stdev ≈ 1.8ms, ≈2.8% of the median), before making
any change:

- **Correctness gate (must pass, non-negotiable):** the after-change command, run against the
  same `week-38.json`, must produce stdout with the same sha256 as `baseline-output.txt`. Any
  difference is a regression regardless of speed, full stop.
- **Faster** = median wall-clock time over 25 timed runs (3 warmup) drops by at least 15% from
  the baseline median, **and** the two runs' noise bands don't overlap
  (`after.median + after.stdev < baseline.median − baseline.stdev`). A change that's faster only
  within the ~2-3ms noise band doesn't count.
- **Regression** = after-median is not clearly below baseline-median by that margin, or is higher,
  or the output hash differs.

## After the change

Two things neither agent did are worth a prompt line. Both tested only the week they were given, though the prompt said the report had slowed down as households grew; the size that motivated the change is the size to measure. And neither tried a week with a household that never came or a pickup for an unknown household, which is where the lesson’s refactor broke. The line: take the baseline on the workload that made it slow, plus these edge cases, and keep the record in the pull request.

How the runs were made and checkedOne run each, recorded as written
  • Both agents received the prompts word for word, in fresh contexts, in the same message, each in a folder seeded with the same report and data. The only differences were the Architecture paragraph and the folder.
  • The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples; the output and timing files they generated are left out, and the architecture agent’s BASELINE.md quotes its numbers. The checker compares every printed line with the original’s on four workloads, then times both on a week four times larger.
  • This is one sample of each prompt, not a measurement of a model. One correct plain refactor does not mean plain prompts preserve behavior; the lesson’s own example is the kind that does not.

09 / Hold it there

Make the snapshot a test, and keep the record next to the code.

A baseline taken once protects one change. Three checks make it protect the next one too.

  1. The tools that already repeat runs

    Go’s go test -bench runs a benchmark until the timing is stable and reports time per operation, and -count repeats it so you can see the spread. In JavaScript, Vitest has a bench mode built on the same idea. Use them for the timing half; they do not check output, so keep the snapshot test beside them.

  2. A snapshot test on the workload and the edge cases

    Store what the report prints for week 38 and for a handful of edge cases, and fail the build when it changes. When the change is intended, the snapshot is updated in the same pull request, where a reviewer sees it. This lesson’s shared cases are that test for all three versions. The checker in section 08 adds the edge cases below, which both agents’ refactors passed without having tried.

    check-runs.mjs
    const edges = {
    	households: [
    		{ id: 'a', name: 'Alder' },
    		{ id: 'b', name: 'Birch' },
    		{ id: 'c', name: 'Cedar' },
    		{ id: 'd', name: 'Dogwood' },
    		{ id: 'e', name: 'Elm' }
    	],
    	pickups: [
    		{ household: 'a', hundredths: 1234 },
    		{ household: 'a', hundredths: 1234 }, // rounding once gives 24.7; rounding each gives 24.6
    		{ household: 'c', hundredths: 1000 },
    		{ household: 'd', hundredths: 1000 }, // a tie with Cedar before Cedar's second pickup
    		{ household: 'e', hundredths: 1245 },
    		{ household: 'c', hundredths: 5 },
    		{ household: 'x', hundredths: 900 } // a pickup for a household not on the list
    	] // Birch never came
    };
  3. A record with the rule in it

    Keep the baseline record in the repository: workload, fingerprint, every sample, runtime, and the rule. The next change is measured against it, by the same rule, instead of against someone’s memory of how fast it used to be.

    bench-typescript.json · before
    {
    	"version": "before",
    	"workload": "week 38: 1200 households, 6000 pickups",
    	"fingerprint": "9a75b00e",
    	"lines": 1200,
    	"work": 7200000,
    	"medianMs": 33.473,
    	"samplesMs": [
    		32.518,
    		32.965,
    		33.781,
    		33.719,
    		33.9,
    		33.141,
    		33.473,
    		34.019,
    		33.671,
    		32.802,
    		32.952,
    		33.413,
    		34.288,
    		33.774,
    		33.309
    	],
    	"runtime": "node v22.21.1 on darwin-arm64",
    	"rule": "same fingerprint; the slowest run after beats the fastest run before"
    }
Build UIs?React and the browser already measure renders. A filter box is where you own the baseline.

Where it already is in your components

React’s <Profiler> “lets you measure rendering performance of a React tree programmatically”, and its onRender callback reports each commit’s actualDuration. It is “disabled in the production build by default” (Profiler), so it measures development builds unless you opt in. In any framework, the browser’s performance.mark and performance.measure put your own timings in the Performance panel. The numbers are there; the habit of keeping them is the part you add.

When you have to own it

The food bank’s household list has a filter box, and someone wants a faster filter. The baseline is the same keystrokes typed before and after, the rows each keystroke shows, and each update’s duration. The rows matter as much as the time: a faster filter that matches differently is a different feature, exactly like the report that dropped the households who did not come.

Collecting render timings: React’s Profiler, and marks and measures around a Svelte update.

ReactAlready in your code
Measured.tsx
import { Profiler, type ProfilerOnRenderCallback, type ReactNode } from 'react';

// React already measures renders. Keep the numbers instead of eyeballing the page: collect
// every commit's duration while you replay the same interaction, before and after a change.
export const commits: { id: string; phase: string; ms: number }[] = [];

const record: ProfilerOnRenderCallback = (id, phase, actualDuration) => {
	commits.push({ id, phase, ms: Math.round(actualDuration * 100) / 100 });
};

export default function Measured({ children }: { children: ReactNode }) {
	return (
		<Profiler id="pickup-list" onRender={record}>
			{children}
		</Profiler>
	);
}

10 / Make the call

Take a baseline when someone will rely on the answer.

A baseline is overkill for a change you would never describe as “faster”, for code nobody depends on yet, and for a script you run once. Take one whenever a change is justified by speed, whenever an agent rewrites something that other people read the output of, and before any change you would want to undo if it turned out worse.

Retake it when the workload grows, when the runtime or machine changes, and when the output is meant to change, so the next comparison starts from the truth.

Take it with you

Explain it without saying “baseline”: “Before changing it, I saved exactly what it printed for this week’s data, timed it fifteen times, and wrote down what I would call faster. After the change it printed the same thing and every run was quicker.” Then find the last change you merged because it felt faster, and write down what you would have measured.

Paste into your next prompt, and fill in the blanks

Before changing <component>, take a baseline and write it to BASELINE.md:
- The workload: <input files or fixtures>, fixed for before and after.
- The behavior: exactly what it outputs on that workload (save it, and a hash).
  Also run <edge cases: empty input, <boundary values>, <odd records>>.
- The measurement: <n> timed runs after <k> warmups, every sample kept.
- The rule, decided now: the output must be identical; "faster" means
  <every run after beats every run before / the median drops by <x>%>.
After the change, measure the same way and report against the rule.
Do not claim a speedup you did not measure.
Connections to follow nextRelated lessons
  • CPU and memory profiling finds where the time goes; this lesson decides whether a change to it helped.
  • Defining success chooses the numbers a user would care about before you measure them.
  • Fitness functions turns a check like the snapshot into one that runs on every change.
  • Spec before code is the same move for new code: write down what must be true before it exists.

Take the report into your editor. Add a column for the last pickup date, and write the snapshot test before you write the column.

Back to architecture →