“Error bar” is a drawing, not a statistical definition.
An interval drawn around a point can represent a range of observed values, standard deviation, standard error, a confidence interval, a prediction interval, or a credible interval. These answer different questions. A standard deviation describes spread among individual observations. A confidence interval describes uncertainty in an estimated parameter under a sampling model. A prediction interval concerns a future observation or outcome.
Suppose a service team counts failed requests during two 30-minute release windows. In each window, “failure” means a completed eligible request returned a 5xx or timed out. The old build had 24 failures among 1,000 eligible requests (2.4%); the candidate had 48 among 1,000 (4.8%). The interval here will be a 95% Wilson score interval for each binomial failure proportion. These are authored examples, not real telemetry.
The chart contract also names the target population: eligible requests served in that region during the specified windows. A percentage without this denominator can hide load changes, filtered requests, or changes in the failure definition.
- Old build
- 24 failures / 1,000 eligible requests = 2.4%
- Candidate
- 48 failures / 1,000 eligible requests = 4.8%
- Measurement unit
- Request-level binary failure outcome
- Interval method
- 95% Wilson score interval for each proportion
For a request failure rate, count events and eligible trials.
For a binomial outcome, each eligible request contributes one trial, and each failure
contributes one success in the counted event. The sample estimate is p = x/n.
Here the old rate is 24/1,000 = 0.024 = 2.4%. The point estimate is exact
arithmetic for the sample; uncertainty arises when we use it to say something about a
broader request population.
This lesson uses the Wilson score interval, a useful interval for a binomial proportion
that behaves better than the simple estimate ± a normal standard error in small samples
and near 0 or 1. With x=24, n=1,000, and z=1.96, it
is approximately 1.61% to 3.56%. For 48/1,000, it is approximately 3.64% to 6.32%. The 1.96 value is the approximate standard-normal critical
value for a two-sided 95% interval.
Put the interpretation beside the graphic.
A useful caption can be short while still being auditable: “Request-level 5xx-or-timeout share among eligible requests, per 30-minute window; points are sample proportions and bars are two-sided 95% Wilson score intervals; old 24/1,000, candidate 48/1,000.” Then specify how requests were sampled or whether the telemetry is a complete census of the window.
Do not round the point and endpoints so aggressively that the difference disappears or precision is exaggerated. Keep enough precision to reproduce the calculation, but avoid implying that the interval endpoints are known with infinite exactness. If the graph is printed without its caption or denominator, it becomes much easier to misread.
- Point
- Sample proportion per build
- Bars
- Two-sided 95% Wilson score interval
- Denominator
- Eligible completed requests, n=1,000 per build
- Window
- Separate 30-minute windows; confirm traffic and region match
Separate intervals on a chart are not the comparison itself.
The visual intervals nearly touch in this illustration. That can attract attention, but “the bars overlap” is not a universal test of whether two rates differ. Separate confidence intervals address uncertainty for each group; a direct comparison needs an interval or test for the difference, using a model appropriate to the data and experiment.
Even a well-calculated interval cannot decide whether a change matters operationally. A rise from 2.4% to 4.8% is a 2.4 percentage-point increase and a 100% relative increase, but impact also depends on total traffic, severity, retry behavior, and the release's user-facing objective. Establish a practical threshold before staring at the plot.
A million requests can still act like only a few independent observations.
Requests in one deployment window may share host load, a dependency incident, a customer cohort, or a single routing change. If failures cluster by host, zone, or minute, a calculation that treats every request as independent can make the interval too narrow. Use the experimental unit and sampling design: perhaps compare deployment blocks, zones, or randomized cohorts rather than pretending the individual request count captures all uncertainty.
Also check whether “eligible” changed, missing telemetry could be related to failure, retries create multiple counted attempts per user operation, or the dashboard selects only successful traces. A proportion describes the outcome as defined and observed. It does not tell you why requests failed, nor does a sample interval contain uncertainty about a misspecified population or biased collection pipeline.
Make the event definition and denominator explicit.
The helper accepts integer event and trial counts and returns percentages. It rejects impossible counts and uses the Wilson score formula. It does not inspect request eligibility, dependence, region, time windows, or causal assignment; those belong in the measurement design. Keep rates as proportions during calculation and convert to percentage only for display.
Both versions use counts, validate the denominator, and label the interval method.
export type Interval = { estimate: number; low: number; high: number; count: number };
/** Wilson score interval for a binomial proportion, shown as percentages. */
export function wilson(successes: number, trials: number, z = 1.96): Interval {
if (
!Number.isInteger(successes) ||
!Number.isInteger(trials) ||
trials < 1 ||
successes < 0 ||
successes > trials
) {
throw new Error('use integer counts with 0 <= successes <= trials and trials >= 1');
}
const p = successes / trials;
const z2 = z * z;
const denominator = 1 + z2 / trials;
const center = (p + z2 / (2 * trials)) / denominator;
const half = (z / denominator) * Math.sqrt((p * (1 - p)) / trials + z2 / (4 * trials * trials));
return {
estimate: p * 100,
low: Math.max(0, center - half) * 100,
high: Math.min(1, center + half) * 100,
count: trials
};
}
package interval
import (
"errors"
"math"
)
type Interval struct {
Estimate, Low, High float64
Count int
}
// Wilson returns a score interval for a binomial proportion, expressed in percentages.
func Wilson(successes, trials int, z float64) (Interval, error) {
if trials < 1 || successes < 0 || successes > trials || z <= 0 || math.IsNaN(z) || math.IsInf(z, 0) {
return Interval{}, errors.New("require trials >= 1, 0 <= successes <= trials, and finite z > 0")
}
p := float64(successes) / float64(trials)
z2 := z * z
denom := 1 + z2/float64(trials)
center := (p + z2/(2*float64(trials))) / denom
half := (z / denom) * math.Sqrt(p*(1-p)/float64(trials)+z2/(4*float64(trials*trials)))
return Interval{p * 100, math.Max(0, center-half) * 100, math.Min(1, center+half) * 100, trials}, nil
}
Use the graph to decide what evidence is missing.
For the meaning and construction of confidence intervals, consult NIST/SEMATECH's confidence interval overview and intervals for a binomial proportion. This lesson uses the Wilson score interval and emphasizes the assumptions behind its interpretation.