← Math in Practice
Concept Samples, variability, and evidence

Confidence, variability, and repeated measurements

One benchmark run is a data point. A decision needs to account for how the measurement changes when you repeat it.

A release candidate appears 4 ms faster in a benchmark. The team has run the test once on each revision. That result could be a real improvement, ordinary run-to-run noise, or a change in the machine or workload. Repeating a measurement does not make the system stable; it lets us see how unstable the evidence is.

The judgment to keep

Report the estimate, the spread of individual measurements, and how uncertain the estimated average is. Then check whether the runs were comparable and whether the change was assigned in a way that supports a causal claim.

TypeScriptGo Sample variability · standard error · confidence intervals · repeated comparisons
01 / Read the apparent change

A faster sample mean is a signal to investigate, not yet a release conclusion.

Suppose a service team runs the same endpoint benchmark 100 times on each build. The measured outcome is request latency in milliseconds. Build A has a sample mean of 216 ms and sample standard deviation of 45 ms. Build B has a mean of 212 ms and sample standard deviation of 40 ms. The measured difference is 212 − 216 = −4 ms.

The sample mean summarizes the observed runs. The sample standard deviation (SD) describes how far individual run measurements typically spread around that sample mean. Neither value alone says whether the 4 ms difference is larger than the variation in the experiment, nor whether latency improved for users.

These are constructed numbers for instruction, not benchmark results. A real report should also name the machine, warm-up, request mix, concurrency, build configuration, collection period, and whether each observation is one request or a run-level summary.

Case file / Two benchmark windowsSame endpoint, two revisions, uncertain difference.
Outcome
Per-run mean latency in milliseconds
Build A
100 runs · mean 216 ms · SD 45 ms
Build B
100 runs · mean 212 ms · SD 40 ms
Observed difference
B − A = −4 ms; negative is faster
02 / Separate spread from precision

Standard deviation and standard error answer different questions.

The sample SD s describes dispersion among measurements. The standard error of the mean describes how much the estimated mean would vary across repeated samples under the model. For independent observations, its estimated value is SE = s / √n.

For Build A, SE = 45 ms / √100 = 4.5 ms. For Build B, SE = 40 ms / √100 = 4 ms. Individual runs vary by tens of milliseconds, while averaging 100 independent runs makes the mean more stable. Increasing sample size can reduce uncertainty in the mean; it does not reduce the variation users experience in individual requests.

That square-root relationship relies on the observations carrying independent information. One hundred back-to-back requests sharing the same cache state, host contention, and deployment may behave more like a small number of independent batches. Count the experimental unit that was independently assigned or reset, not just the rows in a file.

Individual run spread · Build ASD = 45 ms
Mean estimate spread · Build ASE = 45 / √100 = 4.5 ms
03 / Interpret an interval

A 95% confidence interval describes the method across repeated samples.

For a sufficiently large, well-behaved sample, an approximate interval for a mean is x̄ ± 1.96 × SE. Using 1.96 is a teaching approximation to the normal critical value; for small samples or strongly skewed measurements, use an appropriate t-based, robust, or bootstrap method and respect the experiment's design.

For Build A, the approximate interval is 216 ± (1.96 × 4.5), or about [207.2, 224.8] ms. For B it is 212 ± (1.96 × 4), or about [204.2, 219.8] ms. A frequentist 95% procedure means that, under its assumptions, intervals built this way from many comparable samples would cover the fixed population mean about 95% of the time. It does not mean there is a 95% probability that this particular interval contains the mean.

The mean interval communicates uncertainty in the estimated mean. It is not a range expected to contain 95% of individual requests; that is a different object, such as a prediction interval or a percentile range.

04 / Compare repeated measurements

Quantify uncertainty in the difference you actually care about.

When samples are independent, a simple large-sample standard error for the difference of means is SEΔ = √(sA²/nA + sB²/nB). Here that is √(45²/100 + 40²/100) ≈ 6.02 ms. The observed difference B − A is −4 ms, so the approximate 95% interval is −4 ± 1.96 × 6.02, or about [−15.8, 7.8] ms.

This interval includes zero and also includes differences in either direction. The data do not make a clear case that the average changed under this model. That is not proof that the builds are equivalent; a meaningful equivalence claim needs a predeclared tolerance and an interval narrow enough to fit inside it. Nor does overlap of two separate mean intervals serve as a formal comparison test. Analyze the difference directly.

If each run on A is deliberately paired with a run on B using the same host, seed, or workload block, analyze within-pair differences instead. Pairing can remove shared noise, but only if the pairing is real and retained in the data.

Estimated B − A−4.0 ms
Approximate interval for difference[−15.8, 7.8] ms
05 / Challenge the design

A precise comparison can still answer the wrong question.

Before interpreting the interval, ask whether measurements represent the workloads and hardware users encounter; whether runs are independent; whether the same warm-up and load were used; whether another deployment, dependency, or background task changed; and whether the outcome was selected after looking at many possible metrics.

Repeated measurements quantify uncertainty under a data-generating process. They do not eliminate selection bias, instrument error, clock drift, cache effects, autocorrelation, or confounding. A confidence interval around an observational before-and-after change describes the estimated difference, but it cannot alone establish that the code change caused it. Causal attribution needs a design that makes alternative explanations less plausible, such as randomized interleaving across machines or a controlled experiment.

06 / Practice in code

Keep the spread, sample size, and comparison together.

The examples calculate the sample SD with denominator n − 1, the estimated standard error, and a large-sample normal interval. They compare independent samples and label the difference as after minus before. The code helps prevent arithmetic slips; it cannot verify independence, representativeness, or a causal design. For small samples, the fixed 1.96 multiplier is not a substitute for the appropriate t critical value.

Compare the same calculation in TypeScript and Go.

Both examples expose the estimate, sample spread, standard error, and approximation.

TypeScriptRepeated measurements · interval for a mean difference
estimate.ts
export type Summary = {
	count: number;
	mean: number;
	sampleSD: number;
	standardError: number;
	low95: number;
	high95: number;
};

/** Large-sample interval for the mean; observations are assumed independent and representative. */
export function summarize(values: number[]): Summary {
	if (values.length < 2 || values.some((value) => !Number.isFinite(value))) {
		throw new Error('at least two finite measurements are required');
	}
	const mean = values.reduce((sum, value) => sum + value, 0) / values.length;
	const variance =
		values.reduce((sum, value) => sum + (value - mean) ** 2, 0) / (values.length - 1);
	const sampleSD = Math.sqrt(variance);
	const standardError = sampleSD / Math.sqrt(values.length);
	const margin = 1.96 * standardError;
	return {
		count: values.length,
		mean,
		sampleSD,
		standardError,
		low95: mean - margin,
		high95: mean + margin
	};
}

export function compareMeans(before: number[], after: number[]) {
	const a = summarize(before);
	const b = summarize(after);
	const difference = b.mean - a.mean;
	const standardError = Math.sqrt(a.sampleSD ** 2 / a.count + b.sampleSD ** 2 / b.count);
	const margin = 1.96 * standardError;
	return { difference, standardError, low95: difference - margin, high95: difference + margin };
}
GoRepeated measurements · interval for a mean difference
estimate.go
package repeated

import (
	"errors"
	"math"
)

type Summary struct {
	Count                                        int
	Mean, SampleSD, StandardError, Low95, High95 float64
}

// Summarize uses a large-sample normal interval; measurements should be independent and representative.
func Summarize(values []float64) (Summary, error) {
	if len(values) < 2 {
		return Summary{}, errors.New("at least two measurements are required")
	}
	mean := 0.0
	for _, value := range values {
		if math.IsNaN(value) || math.IsInf(value, 0) {
			return Summary{}, errors.New("measurements must be finite")
		}
		mean += value
	}
	mean /= float64(len(values))
	ss := 0.0
	for _, value := range values {
		ss += (value - mean) * (value - mean)
	}
	sd := math.Sqrt(ss / float64(len(values)-1))
	se := sd / math.Sqrt(float64(len(values)))
	margin := 1.96 * se
	return Summary{len(values), mean, sd, se, mean - margin, mean + margin}, nil
}

type Difference struct{ Estimate, StandardError, Low95, High95 float64 }

// CompareIndependent estimates after-minus-before with a Welch-style standard error.
func CompareIndependent(before, after []float64) (Difference, error) {
	a, err := Summarize(before)
	if err != nil {
		return Difference{}, err
	}
	b, err := Summarize(after)
	if err != nil {
		return Difference{}, err
	}
	d := b.Mean - a.Mean
	se := math.Sqrt(a.SampleSD*a.SampleSD/float64(a.Count) + b.SampleSD*b.SampleSD/float64(b.Count))
	m := 1.96 * se
	return Difference{d, se, d - m, d + m}, nil
}
07 / Choose the next experiment

Design a comparison that can separate the plausible explanations.

For the standard error and repeated-sampling interpretation, see the NIST/SEMATECH e-Handbook pages on confidence limits for the mean and what confidence intervals mean. The calculations in this lesson use a normal approximation for illustration; the cited methods explain when a t interval is appropriate.