← Math in Practice
Concept Update a belief with evidence

Conditional probability and Bayes’ rule

An alert changes what you should believe. It does not settle what happened.

At 14:07, the checkout monitor pages: “degradation likely.” The detector catches most real degradations, but the condition is rare. Before waking the database team or rolling back a release, ask what fraction of positive alerts have historically corresponded to the event—and what the alert can tell you about this incident.

The judgment to keep

A detector’s sensitivity answers “how often does it alert when the event is present?” A positive alert asks the reverse question: “how likely is the event given this alert?” Those are different conditional probabilities. Keep the base rate and both denominators visible when you move from one to the other.

TypeScriptGo Base rate · sensitivity · false-positive rate · positive predictive value
01 / Read the alert

The page is evidence. The incident is still a question.

The checkout monitor evaluates one service-region window at a time. An event is present when more than 5% of checkout requests in a five-minute window return a user-visible error. The monitor emits a positive alert when its score crosses a fixed threshold. The on-call engineer sees one positive alert at 14:07.

A positive alert could indicate the defined degradation, or it could be a false positive. A missed alert is also possible. We need the event’s base rate, the detector’s performance when the event is present, and its performance when the event is absent. For this worked example, those rates come from an illustrative set of 10,000 labeled, comparable windows—not from a live service.

Case file / Checkout monitor “Degradation likely” must be interpreted against its calibration population.
Event
>5% user-visible checkout errors in one service-region five-minute window.
Population
10,000 comparable labeled service-region five-minute windows.
Base rate
1% of windows contain the defined degradation.
Question
Given a positive alert, how likely is the event in that window?
Leave with: a precisely defined event, alert, population, and time window. An alert label alone is not a denominator.
02 / Name the rates

Three rates describe three different slices of the evidence.

Let E mean that the defined degradation is present in a window, and let A mean the detector alerts for that window. The base rate is P(E): the fraction of all comparable windows where the event is present. Here it is 100 / 10,000 = 1%.

Sensitivity, also called the true-positive rate, is P(A | E): among only the affected windows, the fraction that alert. It is 90 / 100 = 90%. The false-positive rate is P(A | not E): among only unaffected windows, the fraction that nevertheless alert. It is 99 / 9,900 = 1%. Neither rate is the fraction of alerts that are real.

The reverse conditional, P(E | A), is often called the positive predictive value (PPV) or precision. It is what the page asks for: among positive alerts, how many correspond to the event? The denominator is now all positive alerts, including true and false positives.

01 / Base rateP(E) = 100 / 10,000 = 1%

100 degraded windows among all 10,000 windows.

02 / SensitivityP(A | E) = 90 / 100 = 90%

90 alerts among the 100 degraded windows.

03 / False-positive rateP(A | not E) = 99 / 9,900 = 1%

99 alerts among the 9,900 unaffected windows.

Checkpoint: say each denominator out loud: all windows; affected windows; unaffected windows; then, for PPV, positive-alert windows.
03 / Count the windows

Make the rare event visible before applying the formula.

Start with 10,000 comparable windows. At a 1% base rate, 10,000 × 0.01 = 100 windows contain the event; 10,000 − 100 = 9,900 do not. Sensitivity of 90% means the detector alerts on 100 × 0.90 = 90 affected windows. It misses the other 10. The 1% false-positive rate means it also alerts on 9,900 × 0.01 = 99 unaffected windows.

So the detector emits 90 + 99 = 189 positive alerts in this cohort. Ninety correspond to the event and 99 do not. There are 9,801 true negatives and 10 false negatives. The counts here are whole because the chosen rates and cohort happen to produce whole values; estimated rates in another cohort can produce fractional expected counts.

Illustrative labeled cohort · rows are event status · columns are detector result
Window conditionAlertNo alertTotal
Degradation present90 true positives10 false negatives100
Degradation absent99 false positives9,801 true negatives9,900
Total189 positive alerts9,811 no-alert windows10,000
Sanity check: the alert column must add to 90 + 99 = 189. If the denominator is only 100 affected windows, you are calculating sensitivity, not the probability that an alert is true.
04 / Update the belief

Only 90 of the 189 alerts come from affected windows.

Bayes’ rule combines the event’s base rate with how the detector behaves in each condition: P(E | A) = P(A | E)P(E) / P(A). A positive alert can arise in two ways: the event is present and the detector catches it, or the event is absent and the detector fires by mistake. Therefore P(A) = P(A | E)P(E) + P(A | not E)P(not E).

Substitute the rates: (0.90 × 0.01) / ((0.90 × 0.01) + (0.01 × 0.99)). The numerator is 0.009. The denominator is 0.009 + 0.0099 = 0.0189. So P(E | A) = 0.009 / 0.0189 ≈ 0.476, or about 47.6%. That matches the count: 90 / 189 ≈ 47.6% of positive alerts are true positives in this modeled cohort.

The positive alert raises the probability from 1% before observing the alert to about 47.6% after it. This is a substantial update, but it still leaves the event less likely than not under these assumptions. A posterior is a conditional probability, not a verdict about cause or a command to ignore the page.

01 / True alerts0.90 × 0.01 = 0.009

9 expected true alerts per 1,000 comparable windows.

02 / False alerts0.01 × 0.99 = 0.0099

9.9 expected false alerts per 1,000 windows.

03 / All alerts0.009 + 0.0099 = 0.0189

Positive-alert probability under this model.

04 / Posterior0.009 / 0.0189 ≈ 47.6%

Event probability given a positive alert.

05 / Choose the next check

Look for evidence that separates the plausible explanations.

A positive alert is compatible with a real checkout degradation, a benign shift that confuses the detector, or a telemetry/data-quality fault. “Rollback now” and “ignore it” are both premature conclusions from this alert alone. First check an independent user-outcome measure for the same five-minute window, then compare the detector’s inputs and alert behavior by region, release, and request volume.

If raw checkout errors rise in the same windows and requests share a new release or dependency failure, the event gains support; that still does not uniquely identify the cause. If the detector score rises while independent errors stay flat, inspect feature freshness, missing data, threshold changes, and traffic mix. If one region alone changes, aggregate rates may be hiding a localized condition. Keep these observations separate from their explanations in the incident notes.

Diagnostic hypotheses · predictions before conclusions
Possible explanationWhat it predictsNext evidence to inspect
Checkout degradation is presentIndependent user-visible error counts rise in the same cohort and window.Request outcomes by region, release, status, and total request denominator.
Traffic or feature distribution shiftedAlert inputs change with traffic mix, while the event’s independent definition may not.Raw feature values, request volume, threshold/configuration changes, and score by cohort.
Telemetry is delayed or malformedAlert timing or score disagrees with source timestamps and event counts.Ingestion lag, missing-sample rate, timestamp distribution, and raw event records.
Diagnostic next step: compare the alert with raw checkout outcomes for the same service, region, and five-minute window. Record numerator, denominator, source, and freshness before changing the alert threshold.
06 / Practice in code

Let the calculation carry its inputs and denominator.

These functions take the base rate, sensitivity, and false-positive rate as probabilities between zero and one, plus a cohort size. They return the expected number of true and false alerts and their ratio. With the example inputs—0.01, 0.90, 0.01, and 10_000—the function returns 90 true positives, 99 false positives, and a posterior near 0.47619.

The resulting counts are expectations from rates, not necessarily observed counts; fractional counts can occur for other inputs. The function does not validate whether the rates were measured on representative, correctly labeled windows. That is an empirical question outside the arithmetic.

Compare the posterior calculation in TypeScript and Go.

Both versions make the alert denominator explicit and reject invalid rates.

TypeScriptExpected alert counts and P(event given alert)
alert.ts
export type AlertEstimate = {
	truePositive: number;
	falsePositive: number;
	positiveAlerts: number;
	probabilityEventGivenAlert: number;
};

/**
 * Estimates P(event | alert) from a base rate, sensitivity, and false-positive rate.
 * Inputs are probabilities from 0 to 1. This is a model, not a calibration check.
 */
export function estimateAlertPosterior(
	baseRate: number,
	sensitivity: number,
	falsePositiveRate: number,
	windowCount: number
): AlertEstimate {
	for (const [name, value] of [
		['baseRate', baseRate],
		['sensitivity', sensitivity],
		['falsePositiveRate', falsePositiveRate]
	] as const) {
		if (!Number.isFinite(value) || value < 0 || value > 1) {
			throw new Error(`${name} must be between 0 and 1`);
		}
	}
	if (!Number.isSafeInteger(windowCount) || windowCount < 0) {
		throw new Error('windowCount must be a non-negative safe integer');
	}

	const affectedWindows = windowCount * baseRate;
	const unaffectedWindows = windowCount - affectedWindows;
	const truePositive = affectedWindows * sensitivity;
	const falsePositive = unaffectedWindows * falsePositiveRate;
	const positiveAlerts = truePositive + falsePositive;
	if (positiveAlerts === 0) {
		throw new Error('posterior is undefined when the model predicts no positive alerts');
	}

	return {
		truePositive,
		falsePositive,
		positiveAlerts,
		probabilityEventGivenAlert: truePositive / positiveAlerts
	};
}
GoExpected alert counts and P(event given alert)
alert.go
package mathpractice

import (
	"errors"
	"math"
)

type AlertEstimate struct {
	TruePositive               float64
	FalsePositive              float64
	PositiveAlerts             float64
	ProbabilityEventGivenAlert float64
}

// EstimateAlertPosterior estimates P(event | alert) from a base rate,
// sensitivity, and false-positive rate. Inputs are probabilities in [0, 1].
// This is a model, not a calibration check.
func EstimateAlertPosterior(baseRate, sensitivity, falsePositiveRate float64, windowCount int64) (AlertEstimate, error) {
	for _, value := range []float64{baseRate, sensitivity, falsePositiveRate} {
		if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 || value > 1 {
			return AlertEstimate{}, errors.New("rates must be finite probabilities between 0 and 1")
		}
	}
	if windowCount < 0 || windowCount > (1<<53)-1 {
		return AlertEstimate{}, errors.New("window count must be between 0 and 2^53-1")
	}

	count := float64(windowCount)
	affectedWindows := count * baseRate
	unaffectedWindows := count - affectedWindows
	estimate := AlertEstimate{
		TruePositive:  affectedWindows * sensitivity,
		FalsePositive: unaffectedWindows * falsePositiveRate,
	}
	estimate.PositiveAlerts = estimate.TruePositive + estimate.FalsePositive
	if estimate.PositiveAlerts == 0 {
		return AlertEstimate{}, errors.New("posterior is undefined when the model predicts no positive alerts")
	}
	estimate.ProbabilityEventGivenAlert = estimate.TruePositive / estimate.PositiveAlerts
	return estimate, nil
}
07 / Change the base rate

The same detector means something different in a rarer population.

Keep sensitivity at 90% and false-positive rate at 1%, but imagine the event occurs in only 0.1% of comparable windows. In a new cohort of 10,000 windows, 10 are affected and 9,990 are not. The expected true-positive count is 10 × 0.90 = 9; the expected false-positive count is 9,990 × 0.01 = 99.9. The positive-alert denominator is 9 + 99.9 = 108.9, so the modeled PPV is 9 / 108.9 ≈ 8.3%.

The detector did not change; the population did. This is why a published “90% accurate” claim is incomplete. Accuracy itself mixes several outcomes and depends on how common the event is. For alert interpretation, preserve base rate, sensitivity, false-positive rate, cohort definition, and collection period separately.

In practice, the base rate can vary by service, region, time of day, release phase, and event definition. The detector’s sensitivity and false-positive rate can also shift when its input distribution or threshold changes. Before trusting a posterior, ask when and where each rate was measured, how labels were assigned, and whether alert rates drift across cohorts.

01 / Affected10,000 × 0.001 = 10

Windows with the defined event.

02 / True alerts10 × 0.90 = 9

Expected detections among affected windows.

03 / False alerts9,990 × 0.01 = 99.9

Expected alerts among unaffected windows.

04 / New posterior9 / 108.9 ≈ 8.3%

Expected true alerts divided by all expected alerts.

Question to keep: what is the event, in which population and window, and what are the denominators behind the three rates?