The long-window score looks healthy while the new release is failing now.
At 10:05, 4% of checkout requests served by the canary fail for five minutes. Over the rolling 30-day window, the service has recorded 99.92% success and has not yet exhausted its 99.9% objective. The release dashboard proposes widening traffic because the monthly number remains above target.
This is a deliberately compact incident report. It does not tell us how many users were affected, whether failures cluster on one payment method, whether retries recover the request, or whether the canary and baseline use the same request population. Those are measurements to collect, not details to assume away.
- Objective
- 99.9% good eligible requests over a rolling 30 days.
- Recent canary
- 4% failures in five minutes; eligible request count not yet attached.
- Monthly status
- 99.92% success, aggregated across baseline and canary traffic.
- Decision
- Should the team widen, hold, or roll back while it checks impact?
An objective is a promise over a defined population and time window.
A service level indicator (SLI) is a measurement of a service property users care about. For a request-based availability SLI, define eligible requests and define what counts as good. For example, an eligible checkout request is good if it returns the agreed success response within 800 ms. The exact boundary matters: if a request times out in the browser but eventually succeeds after the client deadline, it may still be bad from the user’s point of view.
An SLO sets a target for that indicator over a window. “99.9% availability” is incomplete without the population, success rule, and window. A team might count eligible requests across all regions over a rolling 30 days. Health checks, internal retries, and administrative traffic should be treated by a documented policy. Excluding a known class can be reasonable if it is outside the user promise; excluding failures after seeing the data turns the indicator into a way to hide impact.
Request-based and time-based SLIs answer different questions. A request-based ratio gives each eligible request equal weight: 9 bad requests out of 10,000 is 0.09%, whether failures occur in one burst or are spread out. A time-based indicator might ask what fraction of one-minute intervals were healthy. It weights periods, not requests; a quiet failing minute can count as much as a busy one. Choose the model that matches the user promise, and do not compare its budget arithmetic to a different model as if they were interchangeable.
A 99.9% objective allows one bad event per thousand eligible requests.
The error budget is the tolerated bad portion implied by an objective.
For a 99.9% request success target, the allowed error fraction is 1 − 0.999 = 0.001, or 0.1%. Over 10 million eligible requests in the defined
window, that is 10,000,000 × 0.001 = 10,000 bad requests. It is not automatically
10,000 seconds of downtime: that conversion belongs to a time-based SLI and its own assumptions.
Suppose the rolling window contains 8,000 bad requests among those 10 million. The measured error rate is 0.08%; 80% of the request budget has been consumed. The objective is still met, but only 2,000 bad requests of budget remain for the rest of that rolling window, assuming the counts and policy stay fixed. The remaining budget is useful operational evidence, not a permission to spend failures casually.
| Quantity | Calculation | Result | Interpretation |
|---|---|---|---|
| Allowed error fraction | 1 − 99.9% | 0.1% | One in 1,000 eligible requests |
| Total allowed bad requests | 10,000,000 × 0.1% | 10,000 | Budget for this population and window |
| Observed bad requests | 8,000 ÷ 10,000,000 | 0.08% | 99.92% good; objective met so far |
| Budget consumed | 8,000 ÷ 10,000 | 80% | 2,000 allowed bad events remain in-window |
Burn rate compares the observed bad-event rate with the rate the objective permits.
For a request-based SLI, the burn rate is observed error rate divided by
the allowed error rate. At a 99.9% target, the allowed error rate is 0.1%. If 4% of
eligible requests are bad in a short window, the burn rate is 4% ÷ 0.1% = 40×. The service is using budget at forty times the
target-compatible rate in that observed population and window.
The ratio is independent of traffic count, so it summarizes severity relative to the objective. It does not report how many people failed: 4 bad requests out of 100 and 40,000 out of 1 million have the same 4% burn rate but very different reach. Keep the numerator, eligible count, route/region/version, and interval beside the ratio.
A related view is budget consumption normalized by elapsed time. If 50% of a 30-day budget
was consumed in 3 days, the simple time-normalized rate is 0.50 ÷ (3/30) = 5×. That calculation assumes the full window’s budget accrues evenly with time. For request
budgets, traffic is rarely even, so use event-rate burn for alerting and interpret
cumulative consumption alongside traffic volume and seasonality. The two calculations are
related views, not identical measurements when volume varies.
A short spike and a sustained regression should not page for the same reason.
A multiwindow alert checks burn over a short window and a longer window at the same time. The short window reacts quickly; the longer one helps reject a tiny transient that would otherwise wake someone without consuming meaningful budget. A common design uses paired windows such as 5 minutes and 1 hour, with a burn threshold selected for the SLO window and response policy. Those particular windows and thresholds are examples, not a universal standard.
For this 30-day objective, one hour is 1/720 of the window. Consuming 2% of the whole
budget in an hour is about 0.02 ÷ (1/720) = 14.4× time-normalized burn. A paired
alert might require both the 5-minute and 1-hour request-rate burns to exceed a policy threshold
near that level. This catches a fast, continuing failure while filtering a one-minute blip.
A team may choose different thresholds based on service criticality, volume, paging tolerance,
and how quickly it can mitigate.
Did the rate just become bad enough to require immediate attention?
Is the damaging rate sustained enough to justify interrupting the on-call engineer?
Include user-facing SLI, eligible counts, affected slice, and a useful runbook action.
Adjust target, request volume, and window rates separately.
The request counts describe one defined population/window. The two rate fields represent independent short and long alert windows.
This calculator handles a request-count SLI. It does not decide whether an alert pages: inspect eligible sample size, affected users, segment, duration, and the team’s response policy. At zero budget (a 100% target), burn is undefined.
With the starting values, 10 million eligible requests at a 99.9% target permit 10,000 bad requests. Four hundred bad requests consume 4% of that allowance. A 4% short-window error rate is 40× burn, while a 0.08% long-window rate is 0.8×. That combination tells a story of a current spike against a still-healthy monthly average: confirm the canary slice and user impact before the monthly dashboard catches up.
Make the eligible count and allowed rate explicit in the calculation.
The helpers calculate request-based allowance, observed rate, budget consumed, and burn for one named window. They reject impossible counts and a target that leaves no error budget. They do not define which requests are eligible, determine whether a failure harms the user, or choose alert thresholds. Those are product and reliability policy decisions.
Both versions preserve integer event counts and expose the target-derived allowed error rate.
export type SLOAssessment = {
target: number;
allowedErrorRate: number;
observedErrorRate: number;
eligibleRequests: number;
badRequests: number;
allowedBadRequests: number;
budgetConsumed: number;
windowBurnRate: number;
};
/** Assess request-based error-budget use for one explicitly defined SLO window. */
export function assessSLO(
target: number,
eligibleRequests: number,
badRequests: number
): SLOAssessment {
if (!Number.isFinite(target) || target <= 0 || target > 1) {
throw new Error('target must be greater than 0 and at most 1');
}
if (!Number.isSafeInteger(eligibleRequests) || eligibleRequests < 0) {
throw new Error('eligibleRequests must be a non-negative safe integer');
}
if (!Number.isSafeInteger(badRequests) || badRequests < 0 || badRequests > eligibleRequests) {
throw new Error('badRequests must be an integer from 0 through eligibleRequests');
}
if (eligibleRequests === 0) throw new Error('eligibleRequests must be positive');
const allowedErrorRate = 1 - target;
if (allowedErrorRate === 0) throw new Error('target must leave a non-zero error budget');
const observedErrorRate = badRequests / eligibleRequests;
const allowedBadRequests = eligibleRequests * allowedErrorRate;
return {
target,
allowedErrorRate,
observedErrorRate,
eligibleRequests,
badRequests,
allowedBadRequests,
budgetConsumed: badRequests / allowedBadRequests,
windowBurnRate: observedErrorRate / allowedErrorRate
};
}
package slo
import "fmt"
type Assessment struct {
Target float64
AllowedErrorRate float64
ObservedErrorRate float64
EligibleRequests int64
BadRequests int64
AllowedBadRequests float64
BudgetConsumed float64
WindowBurnRate float64
}
// AssessRequestSLO measures request-based error-budget use in one defined window.
func AssessRequestSLO(target float64, eligibleRequests, badRequests int64) (Assessment, error) {
if target <= 0 || target > 1 {
return Assessment{}, fmt.Errorf("target must be greater than 0 and at most 1")
}
if eligibleRequests <= 0 {
return Assessment{}, fmt.Errorf("eligibleRequests must be positive")
}
if badRequests < 0 || badRequests > eligibleRequests {
return Assessment{}, fmt.Errorf("badRequests must be from 0 through eligibleRequests")
}
allowed := 1 - target
if allowed <= 0 {
return Assessment{}, fmt.Errorf("target must leave a non-zero error budget")
}
observed := float64(badRequests) / float64(eligibleRequests)
allowedBad := float64(eligibleRequests) * allowed
return Assessment{
Target: target, AllowedErrorRate: allowed, ObservedErrorRate: observed,
EligibleRequests: eligibleRequests, BadRequests: badRequests,
AllowedBadRequests: allowedBad, BudgetConsumed: float64(badRequests) / allowedBad,
WindowBurnRate: observed / allowed,
}, nil
}
Use the budget to structure the decision; use evidence to choose the action.
The compact model is budget = eligible events × (1 − target) and burn = observed bad-event rate ÷ (1 − target). Its denominator carries the
promise. A trustworthy SLO makes the counted population, good outcome, and time window
visible; an actionable alert adds enough context to tell whether users are being hurt and
what the on-call person can do next.
For the original SRE workbook treatment of SLOs and error budgets, see Google’s Implementing SLOs and Alerting on SLOs. Their example thresholds are useful starting points; teams should tune alert policy to their own service, users, and response capacity.