A faster sample mean is a signal to investigate, not yet a release conclusion.
Suppose a service team runs the same endpoint benchmark 100 times on each build. The
measured outcome is request latency in milliseconds. Build A has a sample mean of 216 ms
and sample standard deviation of 45 ms. Build B has a mean of 212 ms and sample standard
deviation of 40 ms. The measured difference is 212 − 216 = −4 ms.
The sample mean summarizes the observed runs. The sample standard deviation (SD) describes how far individual run measurements typically spread around that sample mean. Neither value alone says whether the 4 ms difference is larger than the variation in the experiment, nor whether latency improved for users.
These are constructed numbers for instruction, not benchmark results. A real report should also name the machine, warm-up, request mix, concurrency, build configuration, collection period, and whether each observation is one request or a run-level summary.
- Outcome
- Per-run mean latency in milliseconds
- Build A
- 100 runs · mean 216 ms · SD 45 ms
- Build B
- 100 runs · mean 212 ms · SD 40 ms
- Observed difference
- B − A = −4 ms; negative is faster
Standard deviation and standard error answer different questions.
The sample SD s describes dispersion among measurements. The standard error of the mean describes how much the estimated mean would vary across repeated samples under the model.
For independent observations, its estimated value is SE = s / √n.
For Build A, SE = 45 ms / √100 = 4.5 ms. For Build B, SE = 40 ms / √100 = 4 ms. Individual runs vary by tens of milliseconds, while
averaging 100 independent runs makes the mean more stable. Increasing sample size can
reduce uncertainty in the mean; it does not reduce the variation users experience in
individual requests.
That square-root relationship relies on the observations carrying independent information. One hundred back-to-back requests sharing the same cache state, host contention, and deployment may behave more like a small number of independent batches. Count the experimental unit that was independently assigned or reset, not just the rows in a file.
A 95% confidence interval describes the method across repeated samples.
For a sufficiently large, well-behaved sample, an approximate interval for a mean is x̄ ± 1.96 × SE. Using 1.96 is a teaching approximation to the normal critical value; for small samples
or strongly skewed measurements, use an appropriate t-based, robust, or bootstrap method
and respect the experiment's design.
For Build A, the approximate interval is 216 ± (1.96 × 4.5), or about [207.2, 224.8] ms. For B it is 212 ± (1.96 × 4), or about [204.2, 219.8] ms. A frequentist 95% procedure means that, under its
assumptions, intervals built this way from many comparable samples would cover the fixed
population mean about 95% of the time. It does not mean there is a 95% probability that
this particular interval contains the mean.
The mean interval communicates uncertainty in the estimated mean. It is not a range expected to contain 95% of individual requests; that is a different object, such as a prediction interval or a percentile range.
Quantify uncertainty in the difference you actually care about.
When samples are independent, a simple large-sample standard error for the difference of
means is SEΔ = √(sA²/nA + sB²/nB). Here that is √(45²/100 + 40²/100) ≈ 6.02 ms. The observed difference B − A is −4 ms, so
the approximate 95% interval is −4 ± 1.96 × 6.02, or about [−15.8, 7.8] ms.
This interval includes zero and also includes differences in either direction. The data do not make a clear case that the average changed under this model. That is not proof that the builds are equivalent; a meaningful equivalence claim needs a predeclared tolerance and an interval narrow enough to fit inside it. Nor does overlap of two separate mean intervals serve as a formal comparison test. Analyze the difference directly.
If each run on A is deliberately paired with a run on B using the same host, seed, or workload block, analyze within-pair differences instead. Pairing can remove shared noise, but only if the pairing is real and retained in the data.
A precise comparison can still answer the wrong question.
Before interpreting the interval, ask whether measurements represent the workloads and hardware users encounter; whether runs are independent; whether the same warm-up and load were used; whether another deployment, dependency, or background task changed; and whether the outcome was selected after looking at many possible metrics.
Repeated measurements quantify uncertainty under a data-generating process. They do not eliminate selection bias, instrument error, clock drift, cache effects, autocorrelation, or confounding. A confidence interval around an observational before-and-after change describes the estimated difference, but it cannot alone establish that the code change caused it. Causal attribution needs a design that makes alternative explanations less plausible, such as randomized interleaving across machines or a controlled experiment.
Keep the spread, sample size, and comparison together.
The examples calculate the sample SD with denominator n − 1, the estimated
standard error, and a large-sample normal interval. They compare independent samples and
label the difference as after minus before. The code helps prevent arithmetic slips; it
cannot verify independence, representativeness, or a causal design. For small samples, the
fixed 1.96 multiplier is not a substitute for the appropriate t critical value.
Both examples expose the estimate, sample spread, standard error, and approximation.
export type Summary = {
count: number;
mean: number;
sampleSD: number;
standardError: number;
low95: number;
high95: number;
};
/** Large-sample interval for the mean; observations are assumed independent and representative. */
export function summarize(values: number[]): Summary {
if (values.length < 2 || values.some((value) => !Number.isFinite(value))) {
throw new Error('at least two finite measurements are required');
}
const mean = values.reduce((sum, value) => sum + value, 0) / values.length;
const variance =
values.reduce((sum, value) => sum + (value - mean) ** 2, 0) / (values.length - 1);
const sampleSD = Math.sqrt(variance);
const standardError = sampleSD / Math.sqrt(values.length);
const margin = 1.96 * standardError;
return {
count: values.length,
mean,
sampleSD,
standardError,
low95: mean - margin,
high95: mean + margin
};
}
export function compareMeans(before: number[], after: number[]) {
const a = summarize(before);
const b = summarize(after);
const difference = b.mean - a.mean;
const standardError = Math.sqrt(a.sampleSD ** 2 / a.count + b.sampleSD ** 2 / b.count);
const margin = 1.96 * standardError;
return { difference, standardError, low95: difference - margin, high95: difference + margin };
}
package repeated
import (
"errors"
"math"
)
type Summary struct {
Count int
Mean, SampleSD, StandardError, Low95, High95 float64
}
// Summarize uses a large-sample normal interval; measurements should be independent and representative.
func Summarize(values []float64) (Summary, error) {
if len(values) < 2 {
return Summary{}, errors.New("at least two measurements are required")
}
mean := 0.0
for _, value := range values {
if math.IsNaN(value) || math.IsInf(value, 0) {
return Summary{}, errors.New("measurements must be finite")
}
mean += value
}
mean /= float64(len(values))
ss := 0.0
for _, value := range values {
ss += (value - mean) * (value - mean)
}
sd := math.Sqrt(ss / float64(len(values)-1))
se := sd / math.Sqrt(float64(len(values)))
margin := 1.96 * se
return Summary{len(values), mean, sd, se, mean - margin, mean + margin}, nil
}
type Difference struct{ Estimate, StandardError, Low95, High95 float64 }
// CompareIndependent estimates after-minus-before with a Welch-style standard error.
func CompareIndependent(before, after []float64) (Difference, error) {
a, err := Summarize(before)
if err != nil {
return Difference{}, err
}
b, err := Summarize(after)
if err != nil {
return Difference{}, err
}
d := b.Mean - a.Mean
se := math.Sqrt(a.SampleSD*a.SampleSD/float64(a.Count) + b.SampleSD*b.SampleSD/float64(b.Count))
m := 1.96 * se
return Difference{d, se, d - m, d + m}, nil
}
Design a comparison that can separate the plausible explanations.
For the standard error and repeated-sampling interpretation, see the NIST/SEMATECH e-Handbook pages on confidence limits for the mean and what confidence intervals mean. The calculations in this lesson use a normal approximation for illustration; the cited methods explain when a t interval is appropriate.