← Math in Practice
Concept Math in engineering decisions

SLOs, error budgets, and burn rate

A reliability target becomes useful when the team can say what it measures and what it changes.

A payments team is deciding whether to widen a checkout release. In the last five minutes, the new version returned errors for 4% of eligible requests. The 30-day dashboard still says the service is 99.92% available against a 99.9% objective. One number describes recent damage; the other averages a larger window. Neither number alone answers whether to keep sending users to the release.

The judgment to keep

Name the user-relevant event, the good outcome, the eligible population, and the time window before calculating a budget. Read short and long windows together, then choose an action that matches observed impact and the cost of being wrong.

TypeScriptGo SLI population · objective · error budget · burn rate · multiwindow alerts
01 / Read the release report

The long-window score looks healthy while the new release is failing now.

At 10:05, 4% of checkout requests served by the canary fail for five minutes. Over the rolling 30-day window, the service has recorded 99.92% success and has not yet exhausted its 99.9% objective. The release dashboard proposes widening traffic because the monthly number remains above target.

This is a deliberately compact incident report. It does not tell us how many users were affected, whether failures cluster on one payment method, whether retries recover the request, or whether the canary and baseline use the same request population. Those are measurements to collect, not details to assume away.

Case file / Checkout canaryA monthly pass mark can hide a sharp local regression.
Objective
99.9% good eligible requests over a rolling 30 days.
Recent canary
4% failures in five minutes; eligible request count not yet attached.
Monthly status
99.92% success, aggregated across baseline and canary traffic.
Decision
Should the team widen, hold, or roll back while it checks impact?
Start with the population: what operation is being measured, who can be affected, and which outcomes count as good?
02 / Define what counts

An objective is a promise over a defined population and time window.

A service level indicator (SLI) is a measurement of a service property users care about. For a request-based availability SLI, define eligible requests and define what counts as good. For example, an eligible checkout request is good if it returns the agreed success response within 800 ms. The exact boundary matters: if a request times out in the browser but eventually succeeds after the client deadline, it may still be bad from the user’s point of view.

An SLO sets a target for that indicator over a window. “99.9% availability” is incomplete without the population, success rule, and window. A team might count eligible requests across all regions over a rolling 30 days. Health checks, internal retries, and administrative traffic should be treated by a documented policy. Excluding a known class can be reasonable if it is outside the user promise; excluding failures after seeing the data turns the indicator into a way to hide impact.

Request-based and time-based SLIs answer different questions. A request-based ratio gives each eligible request equal weight: 9 bad requests out of 10,000 is 0.09%, whether failures occur in one burst or are spread out. A time-based indicator might ask what fraction of one-minute intervals were healthy. It weights periods, not requests; a quiet failing minute can count as much as a busy one. Choose the model that matches the user promise, and do not compare its budget arithmetic to a different model as if they were interchangeable.

Request-based SLIgood eligible requests ÷ all eligible requests
Objective over the windowSLI ≥ target
03 / Calculate the budget

A 99.9% objective allows one bad event per thousand eligible requests.

The error budget is the tolerated bad portion implied by an objective. For a 99.9% request success target, the allowed error fraction is 1 − 0.999 = 0.001, or 0.1%. Over 10 million eligible requests in the defined window, that is 10,000,000 × 0.001 = 10,000 bad requests. It is not automatically 10,000 seconds of downtime: that conversion belongs to a time-based SLI and its own assumptions.

Suppose the rolling window contains 8,000 bad requests among those 10 million. The measured error rate is 0.08%; 80% of the request budget has been consumed. The objective is still met, but only 2,000 bad requests of budget remain for the rest of that rolling window, assuming the counts and policy stay fixed. The remaining budget is useful operational evidence, not a permission to spend failures casually.

Illustrative rolling 30-day request budget
QuantityCalculationResultInterpretation
Allowed error fraction1 − 99.9%0.1%One in 1,000 eligible requests
Total allowed bad requests10,000,000 × 0.1%10,000Budget for this population and window
Observed bad requests8,000 ÷ 10,000,0000.08%99.92% good; objective met so far
Budget consumed8,000 ÷ 10,00080%2,000 allowed bad events remain in-window
04 / Interpret the burn rate

Burn rate compares the observed bad-event rate with the rate the objective permits.

For a request-based SLI, the burn rate is observed error rate divided by the allowed error rate. At a 99.9% target, the allowed error rate is 0.1%. If 4% of eligible requests are bad in a short window, the burn rate is 4% ÷ 0.1% = 40×. The service is using budget at forty times the target-compatible rate in that observed population and window.

The ratio is independent of traffic count, so it summarizes severity relative to the objective. It does not report how many people failed: 4 bad requests out of 100 and 40,000 out of 1 million have the same 4% burn rate but very different reach. Keep the numerator, eligible count, route/region/version, and interval beside the ratio.

A related view is budget consumption normalized by elapsed time. If 50% of a 30-day budget was consumed in 3 days, the simple time-normalized rate is 0.50 ÷ (3/30) = 5×. That calculation assumes the full window’s budget accrues evenly with time. For request budgets, traffic is rarely even, so use event-rate burn for alerting and interpret cumulative consumption alongside traffic volume and seasonality. The two calculations are related views, not identical measurements when volume varies.

Observed request-rate burnbad / eligible ÷ (1 − target)
Time-normalized budget usebudget consumed ÷ window elapsed
40× is a rate, not a forecast: it does not promise the budget will be exhausted in exactly a particular number of hours. Traffic, recovery, rolling-window boundaries, and measurement lag all matter.
05 / Read both alert windows

A short spike and a sustained regression should not page for the same reason.

A multiwindow alert checks burn over a short window and a longer window at the same time. The short window reacts quickly; the longer one helps reject a tiny transient that would otherwise wake someone without consuming meaningful budget. A common design uses paired windows such as 5 minutes and 1 hour, with a burn threshold selected for the SLO window and response policy. Those particular windows and thresholds are examples, not a universal standard.

For this 30-day objective, one hour is 1/720 of the window. Consuming 2% of the whole budget in an hour is about 0.02 ÷ (1/720) = 14.4× time-normalized burn. A paired alert might require both the 5-minute and 1-hour request-rate burns to exceed a policy threshold near that level. This catches a fast, continuing failure while filtering a one-minute blip. A team may choose different thresholds based on service criticality, volume, paging tolerance, and how quickly it can mitigate.

Short window / 5 minFast signal

Did the rate just become bad enough to require immediate attention?

Long window / 1 hourPersistence check

Is the damaging rate sustained enough to justify interrupting the on-call engineer?

Both exceed policyPage with context

Include user-facing SLI, eligible counts, affected slice, and a useful runbook action.

06 / Change the measurements

Adjust target, request volume, and window rates separately.

Request SLO labAllowed error = 1 − target   ·   Burn = observed rate ÷ allowed rate

The request counts describe one defined population/window. The two rate fields represent independent short and long alert windows.

Allowed error rate0.100%100% − 99.900% target
Budget allowance10,000 bad10,000,000 eligible × allowed fraction
Budget consumed4.0%400 observed bad requests in this window
Request-window burn0.0×0.004% observed ÷ 0.100% allowed
Short-window burn40.0×4.00% observed
Long-window burn0.8×0.08% observed

This calculator handles a request-count SLI. It does not decide whether an alert pages: inspect eligible sample size, affected users, segment, duration, and the team’s response policy. At zero budget (a 100% target), burn is undefined.

With the starting values, 10 million eligible requests at a 99.9% target permit 10,000 bad requests. Four hundred bad requests consume 4% of that allowance. A 4% short-window error rate is 40× burn, while a 0.08% long-window rate is 0.8×. That combination tells a story of a current spike against a still-healthy monthly average: confirm the canary slice and user impact before the monthly dashboard catches up.

07 / Practice in code

Make the eligible count and allowed rate explicit in the calculation.

The helpers calculate request-based allowance, observed rate, budget consumed, and burn for one named window. They reject impossible counts and a target that leaves no error budget. They do not define which requests are eligible, determine whether a failure harms the user, or choose alert thresholds. Those are product and reliability policy decisions.

Compare the same SLO calculation in TypeScript and Go.

Both versions preserve integer event counts and expose the target-derived allowed error rate.

TypeScriptRequest SLO assessment · target and event counts
slo.ts · request error budget and burn
export type SLOAssessment = {
	target: number;
	allowedErrorRate: number;
	observedErrorRate: number;
	eligibleRequests: number;
	badRequests: number;
	allowedBadRequests: number;
	budgetConsumed: number;
	windowBurnRate: number;
};

/** Assess request-based error-budget use for one explicitly defined SLO window. */
export function assessSLO(
	target: number,
	eligibleRequests: number,
	badRequests: number
): SLOAssessment {
	if (!Number.isFinite(target) || target <= 0 || target > 1) {
		throw new Error('target must be greater than 0 and at most 1');
	}
	if (!Number.isSafeInteger(eligibleRequests) || eligibleRequests < 0) {
		throw new Error('eligibleRequests must be a non-negative safe integer');
	}
	if (!Number.isSafeInteger(badRequests) || badRequests < 0 || badRequests > eligibleRequests) {
		throw new Error('badRequests must be an integer from 0 through eligibleRequests');
	}
	if (eligibleRequests === 0) throw new Error('eligibleRequests must be positive');

	const allowedErrorRate = 1 - target;
	if (allowedErrorRate === 0) throw new Error('target must leave a non-zero error budget');
	const observedErrorRate = badRequests / eligibleRequests;
	const allowedBadRequests = eligibleRequests * allowedErrorRate;
	return {
		target,
		allowedErrorRate,
		observedErrorRate,
		eligibleRequests,
		badRequests,
		allowedBadRequests,
		budgetConsumed: badRequests / allowedBadRequests,
		windowBurnRate: observedErrorRate / allowedErrorRate
	};
}
GoRequest SLO assessment · target and event counts
slo.go · request error budget and burn
package slo

import "fmt"

type Assessment struct {
	Target             float64
	AllowedErrorRate   float64
	ObservedErrorRate  float64
	EligibleRequests   int64
	BadRequests        int64
	AllowedBadRequests float64
	BudgetConsumed     float64
	WindowBurnRate     float64
}

// AssessRequestSLO measures request-based error-budget use in one defined window.
func AssessRequestSLO(target float64, eligibleRequests, badRequests int64) (Assessment, error) {
	if target <= 0 || target > 1 {
		return Assessment{}, fmt.Errorf("target must be greater than 0 and at most 1")
	}
	if eligibleRequests <= 0 {
		return Assessment{}, fmt.Errorf("eligibleRequests must be positive")
	}
	if badRequests < 0 || badRequests > eligibleRequests {
		return Assessment{}, fmt.Errorf("badRequests must be from 0 through eligibleRequests")
	}
	allowed := 1 - target
	if allowed <= 0 {
		return Assessment{}, fmt.Errorf("target must leave a non-zero error budget")
	}
	observed := float64(badRequests) / float64(eligibleRequests)
	allowedBad := float64(eligibleRequests) * allowed
	return Assessment{
		Target: target, AllowedErrorRate: allowed, ObservedErrorRate: observed,
		EligibleRequests: eligibleRequests, BadRequests: badRequests,
		AllowedBadRequests: allowedBad, BudgetConsumed: float64(badRequests) / allowedBad,
		WindowBurnRate: observed / allowed,
	}, nil
}
08 / Make the release call

Use the budget to structure the decision; use evidence to choose the action.

The compact model is budget = eligible events × (1 − target) and burn = observed bad-event rate ÷ (1 − target). Its denominator carries the promise. A trustworthy SLO makes the counted population, good outcome, and time window visible; an actionable alert adds enough context to tell whether users are being hurt and what the on-call person can do next.

For the original SRE workbook treatment of SLOs and error budgets, see Google’s Implementing SLOs and Alerting on SLOs. Their example thresholds are useful starting points; teams should tune alert policy to their own service, users, and response capacity.