← Math in Practice
Concept Reliability, probability, and evidence

Availability and failure probability

Redundancy changes a probability only when the failures can actually differ.

A checkout path now depends on an API, an identity service, and a database. A proposal says that two API replicas will make the path “four nines.” Before putting that number in a review, ask what is being counted, over which time window, and which outages could take both replicas down together.

The judgment to keep

Calculate only after naming the service boundary and window. Multiply required dependencies only under stated independence assumptions, and treat redundancy as conditional on genuinely separate failure modes.

TypeScriptGo Time-based availability · series paths · parallel replicas · common-mode failure
01 / Read the reliability claim

A service path fails when any required part is unavailable.

In this illustrative release review, the team has monthly time-based availability estimates of 99.95% for the checkout API, 99.90% for identity, and 99.95% for the database. The path needs all three to serve a checkout. The numbers are authored examples, not production observations.

Availability describes the fraction of a defined interval when a specified service boundary met its availability rule. The rule might be “health checks pass” or “eligible requests meet the service objective”; those are not interchangeable. Here the illustrative figures are monthly time fractions over one 30-day calendar month, or 43,200 minutes. They are not probabilities that a particular request succeeds.

Case file / Checkout reliability reviewThree component estimates, one user-facing path.
API
99.95% time availability during the stated month.
Identity
99.90% time availability during the same interval.
Database
99.95% time availability during the same interval.
Question
What can multiplication estimate, and what assumption would make it wrong?
02 / Calculate a required path

For a series path, every required component must be up.

If a request requires components A, B, and C, the path is available only when all three are available at the same time. If their states are independent at the time scale being measured, then Apath = AA × AB × AC. This is a series system: failure of any required component breaks the path.

Substitute fractions, not whole-number percentages: 0.9995 × 0.9990 × 0.9995 = 0.99800124975, or 99.800124975%. For a 30-day interval, expected unavailable time under the stationary model is (1 − 0.99800124975) × 43,200 min ≈ 86.35 min. This is a model-based conversion of a time fraction; it does not say when the downtime occurs.

Multiplying measured marginal availabilities does not generally give the actual joint availability. It gives an estimate when independence is credible. If component states are correlated, use aligned incident timelines to measure when the whole path was unavailable, or build a model that represents the shared causes.

Component fractions0.9995 × 0.9990 × 0.9995
Path estimate0.99800124975 = 99.8001%
Window30 × 24 × 60 = 43,200 min
Unavailable time estimate(1 − 0.99800124975) × 43,200 ≈ 86.35 min
03 / Estimate redundant replicas

Either replica can serve the work, if either can fail on its own.

For two equivalent replicas, the redundant group is down only when both replicas are down. If their failures are independent and each has availability A, then Aparallel = 1 − (1 − A)². With two 99.9% replicas, that is 1 − (0.001 × 0.001) = 0.999999, or 99.9999% for the replica group.

For n independent interchangeable replicas, estimate 1 − (1 − A)ⁿ. “Interchangeable” matters: replicas must have compatible data, capacity, routing, and permissions, and the load balancer must direct work to a healthy one. The equation does not include detection delay, failover errors, overload of the survivor, or recovery time unless those effects are represented in the input measurements.

One replica unavailable1 − 0.999 = 0.001
Both unavailable, if independent0.001 × 0.001 = 0.000001
At least one available1 − 0.000001 = 0.999999 = 99.9999%
04 / Find the shared failure domain

Two copies do little for an outage that reaches both copies.

Suppose the two API replicas each measure 99.9% availability, but both sit behind one 99.9%-available regional network dependency. The independence-only replica calculation says 99.9999%. When the network dependency is down, though, neither replica is reachable. If its state is independent of replica failures, include it in series: 0.999 × 0.999999 = 0.998999001, or about 99.8999% for this simplified path.

The shared dependency now dominates; adding more replicas in the same failure domain cannot raise path availability above that dependency’s 99.9% availability. If the network failures also correlate with replica failures, even that product is not justified. A measured joint availability or a dependency model with common causes is needed.

Look for shared regions and zones, power and network paths, DNS, identity and secrets, storage, control planes, deploy pipelines, and human actions. Redundancy is a topology and operating property as well as a replica count.

05 / Change the assumptions

Put the shared dependency back into the estimate.

Change the values to compare a three-service series path with a redundant replica group. The window is held at 30 days (43,200 minutes). The calculator treats the entered percentages as time availability estimates; it cannot decide whether the definitions, evidence, or independence assumptions are valid.

Required checkout path · series
Replica group · interchangeable copies
Series estimate · independent component states99.800125%

Estimated unavailable time: 86.35 minutes per 30-day window.

Independent replica-only estimate99.999900%

Assumes independent replica states and equivalent capacity.

With shared dependency in series99.899900%

Formula: shared availability × independent replica-group availability.

These are time-fraction estimates, not per-request success probabilities. For simplicity, the shared dependency is assumed independent of replica failures; real common causes can violate that assumption. Confirm availability rules and failure domains against incident timelines and end-to-end SLI data.

06 / Practice in code

Make the assumptions visible at the function boundary.

The helpers take availability percentages and an explicit window in minutes. The series helper multiplies component fractions; the redundant helper combines interchangeable replica availability with an independent shared dependency. Input validation catches values outside 0–100% and invalid replica counts, but code cannot validate whether the operational assumptions are true.

Compare the same model in TypeScript and Go.

Both examples preserve the time unit and document the independence assumptions.

TypeScriptAvailability estimates · independence assumptions documented
availability.ts
export type AvailabilityEstimate = {
	availabilityPercent: number;
	downtimeMinutes: number;
};

function validateAvailability(value: number, label: string): void {
	if (!Number.isFinite(value) || value < 0 || value > 100) {
		throw new Error(`${label} must be between 0 and 100 percent`);
	}
}

/** Estimates a series path where every component must be available.
 * This product model assumes component states are independent in the stated window.
 */
export function estimateSeriesAvailability(
	componentAvailabilityPercent: number[],
	windowMinutes: number
): AvailabilityEstimate {
	if (componentAvailabilityPercent.length === 0) {
		throw new Error('at least one component is required');
	}
	if (!Number.isFinite(windowMinutes) || windowMinutes < 0) {
		throw new Error('windowMinutes must be finite and non-negative');
	}

	const availability = componentAvailabilityPercent.reduce((product, component, index) => {
		validateAvailability(component, `component ${index + 1}`);
		return product * (component / 100);
	}, 1);

	return {
		availabilityPercent: availability * 100,
		downtimeMinutes: (1 - availability) * windowMinutes
	};
}

/** Estimates interchangeable replicas plus one independent shared dependency.
 * The replica failures must be independent conditional on the shared dependency being up.
 */
export function estimateRedundantAvailability(
	replicaAvailabilityPercent: number,
	replicaCount: number,
	sharedDependencyAvailabilityPercent: number
): number {
	validateAvailability(replicaAvailabilityPercent, 'replica availability');
	validateAvailability(sharedDependencyAvailabilityPercent, 'shared dependency availability');
	if (!Number.isSafeInteger(replicaCount) || replicaCount < 1) {
		throw new Error('replicaCount must be a positive safe integer');
	}

	const replicaAvailability = replicaAvailabilityPercent / 100;
	const sharedAvailability = sharedDependencyAvailabilityPercent / 100;
	const allReplicasUnavailable = (1 - replicaAvailability) ** replicaCount;
	return sharedAvailability * (1 - allReplicasUnavailable) * 100;
}
GoAvailability estimates · independence assumptions documented
availability.go
package availability

import (
	"errors"
	"fmt"
	"math"
)

type Estimate struct {
	AvailabilityPercent float64
	DowntimeMinutes     float64
}

func validateAvailability(value float64, label string) error {
	if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 || value > 100 {
		return fmt.Errorf("%s must be between 0 and 100 percent", label)
	}
	return nil
}

// Series estimates a path where every component must be available.
// The product model assumes independent component states in the stated window.
func Series(componentsPercent []float64, windowMinutes float64) (Estimate, error) {
	if len(componentsPercent) == 0 {
		return Estimate{}, errors.New("at least one component is required")
	}
	if math.IsNaN(windowMinutes) || math.IsInf(windowMinutes, 0) || windowMinutes < 0 {
		return Estimate{}, errors.New("windowMinutes must be finite and non-negative")
	}

	availability := 1.0
	for index, component := range componentsPercent {
		if err := validateAvailability(component, fmt.Sprintf("component %d", index+1)); err != nil {
			return Estimate{}, err
		}
		availability *= component / 100
	}

	return Estimate{
		AvailabilityPercent: availability * 100,
		DowntimeMinutes:     (1 - availability) * windowMinutes,
	}, nil
}

// Redundant estimates interchangeable replicas plus one independent shared dependency.
// Replica failures must be independent conditional on the shared dependency being up.
func Redundant(replicaPercent float64, replicaCount int, sharedDependencyPercent float64) (float64, error) {
	if err := validateAvailability(replicaPercent, "replica availability"); err != nil {
		return 0, err
	}
	if err := validateAvailability(sharedDependencyPercent, "shared dependency availability"); err != nil {
		return 0, err
	}
	if replicaCount < 1 {
		return 0, errors.New("replicaCount must be at least 1")
	}

	replicaAvailability := replicaPercent / 100
	sharedAvailability := sharedDependencyPercent / 100
	allReplicasUnavailable := math.Pow(1-replicaAvailability, float64(replicaCount))
	return sharedAvailability * (1 - allReplicasUnavailable) * 100, nil
}
07 / Choose the next measurement

Use the estimate to ask better questions of the real system.

Availability and error budget math is a way to reason about tolerated service risk, not a substitute for observing user outcomes. See the Google SRE Book chapters on embracing risk and the availability table for the relationship between SLOs and time budgets.

For the architectural limit on simple redundancy, see the Google SRE Book’s discussion of production environments and failure domains. The lesson’s formulas are simplified teaching models; a production claim should be based on a defined SLI and evidence from the actual system.