“Faster” depends on the work you mean.
An inbox stores messages in 97 threads. One implementation scans every message for each thread request. The candidate builds an index once and then reads the matching thread. Both should return the same message IDs in the same order.
For one thread request, index setup is part of the user-visible question. For a batch of 24 requests, the same setup can be amortized across repeated replay. Those are two valid questions with different workloads; they should not share a headline number.
One catalog snapshot · one query list · two equivalent implementations.
- Workload
- Choose one request or a repeated batch against the same generated inbox.
- Timed region
- Session setup plus replay, with data generation and output comparison outside the clock.
- Reported evidence
- Median and range from repeated samples, plus visits, returned IDs, and index entries.
- Decision
- Only claim what this workload and measurement plan can distinguish.
From “is it faster?” to defensible evidence.
Each step fixes one source of ambiguity. If the result changes when you change the workload, that is a finding about the question, not a reason to hide the earlier run.
- 01 Name the question
Choose latency, throughput, or setup trade-off.
A one-request lookup and a batch of repeated lookups exercise different costs. Write the user-shaped workload before choosing an input size.
Leave withA workload you can replay unchanged for both implementations. - 02 Check equivalence
Prove the outputs still match.
Compare returned IDs, order, and count outside the timed region. A fast wrong answer is not a performance win.
Leave withA correctness result that gates interpretation of the timings. - 03 Set the boundary
Decide what setup belongs in the claim.
If users pay for index construction during a request, include it. If it is a startup cost paid once, say so and measure replay separately rather than quietly dropping it.
Leave withA timed region whose inclusion rules are written down. - 04 Warm and repeat
Let startup settle, then show variability.
Warm both paths before reporting. Alternate order where practical and report a median plus range from several samples; do not select the fastest run.
Leave withA center and a spread, not a false precision. - 05 Interpret
Make a conditional claim.
Report the machine, runtime, workload, sample plan, and output check. A result supports “for this batch on this process,” not “always faster.”
Leave withA reproducible comparison and the next workload worth checking.
Run a measurement plan, then read its assumptions.
The lab runs the displayed TypeScript implementation. Change the question between one request and a batch, or compare a first run with a warmed repeated plan. The output check and operation counts remain visible so the timing has context.
What question does this run answer?
The displayed TypeScript source runs here. The browser is the measurement environment, so treat timings as observations and use the operation counts to explain the work.
Prediction: a 24-request question gives setup work a chance to amortize.
Two warmup rounds, then five reported samples with alternating order.
Choose a workload, predict what it measures, then run the comparison.
Expect output equality, a median, a range, and counts that describe the chosen workload.
Browser timings include this process’s runtime and instrumentation. They are not a cross-machine benchmark, a throughput guarantee, or an allocated-byte measurement. Change one workload or plan at a time.
Why the first run is not automatically the truthStartup, compilation, caches, and scheduling
A first run can include runtime startup, compilation, cache population, or an unrelated scheduling pause. It is still useful when the product question is cold-start latency. It is not a substitute for a repeated steady-state plan.
Hold the workload steady. Change the language.
The generated inbox and query list stay outside the timed region.
export function makeMessages(count: number): Message[] {
return Array.from({ length: count }, (_, sequence) => ({
id: `message-${sequence}`,
threadId: `thread-${sequence % THREADS}`,
sequence,
body: `message ${sequence} in thread ${sequence % THREADS}`
}));
}
export function makeQueries(workload: Workload, queryCount: number): string[] {
const count = workload === 'single' ? 1 : queryCount;
return Array.from({ length: count }, (_, index) => `thread-${(index * 13 + 7) % THREADS}`);
}
export function median(values: number[]): number {
const sorted = [...values].sort((a, b) => a - b);
return sorted[Math.floor(sorted.length / 2)] ?? 0;
} func makeMessages(count int) []Message {
messages := make([]Message, count)
for sequence := 0; sequence < count; sequence++ {
thread := sequence % threads
messages[sequence] = Message{
ID: fmt.Sprintf("message-%d", sequence),
ThreadID: fmt.Sprintf("thread-%d", thread),
Sequence: sequence,
Body: fmt.Sprintf("message %d in thread %d", sequence, thread),
}
}
return messages
}
func makeQueries(workload Workload, queryCount int) []string {
count := queryCount
if workload == Single {
count = 1
}
queries := make([]string, count)
for index := range queries {
queries[index] = fmt.Sprintf("thread-%d", (index*13+7)%threads)
}
return queries
} The input generator and query list are deterministic, but they stay outside the clock. The equivalence check compares the outputs of the first timed samples after the clock stops, so nothing warms either path before a first-run sample. The repeated plan alternates order, excludes two warmup rounds from the report, and retains five observed samples. The call-site summary refuses to interpret unequal outputs.
Copy the complete examplesStandard library only
These files are runnable without a benchmark framework. A framework can add statistics and isolation, but it cannot choose the workload or decide whether setup belongs in your claim.
TypeScriptnode --experimental-strip-types inbox.ts
Gogo run inbox.go
A benchmark is evidence about a setup, not a universal ranking.
This example establishes that the two paths return equivalent IDs for the selected generated workload and shows how setup, replay, and sample spread behave in the current process. It does not establish production latency, memory allocation, tail latency, multi-user contention, garbage-collection behavior under load, or performance on another runtime.
For those questions, use a production-shaped dataset, a load test, a memory profiler, or a tracing system. Keep the benchmark small enough to run repeatedly, but do not confuse a toy workload with a representative one.
Build UIs?See where this shows up in your components.
Frameworks help with mechanics, not judgment.
A JavaScript benchmark library and Go’s testing.B can manage repetitions and statistics.
You still own the fixture, setup boundary, equivalence check, and claim. A framework default
is not automatically the right experiment.
Choose the plan that preserves the question.
You want to say whether the indexed path helps a typical batch. One path currently receives fewer queries than the other. Select the plan that gives you comparable evidence.
You want to claim the index helps a typical batch.
Which measurement plan gives you evidence you can interpret?
Make the next comparison cheaper.
Record the workload before the number. Leave a note that says why you measured, what you timed and how, what the number cannot speak for, what you do when the evidence is unclear, and when to measure again.
- Why
- The index answers the same inbox query as the scan but costs setup, so “faster” depends on how many queries share that setup.
- What
- Both paths get the same generated messages and queries. Outputs are compared after the clock stops. The repeated plan alternates order, drops two warmup rounds, and reports the median and range of five samples.
- Constraint
- The numbers describe this runtime, this machine, and this generated workload. Tail latency, memory, contention, and other input mixes are separate questions.
- Fallback
- Unequal outputs mean no timing claim at all. When the ranges overlap more than the medians differ, report no clear difference instead of picking a run.
- Reconsider when
- The batch size, message count, or query mix changes, the runtime is upgraded, or the index starts being reused across requests.
A procedure note to adapt to your own workload. Nothing here is saved to an account.
Explore more concepts & practices →