← Math in Practice
Concept Measure what a classifier catches and flags

Precision, recall, and the confusion matrix

A threshold moves work between missed events and unnecessary investigations.

The account-protection team has a queue full of sign-ins marked “suspicious.” A lower score threshold might surface more takeovers before they succeed, but it also sends more legitimate customers for review. A stricter threshold protects analyst capacity while leaving more real attacks unflagged. Before someone calls the model “90% accurate,” count what it got right and wrong, and say which outcome matters for the next action.

The judgment to keep

Precision asks what share of flags are real events; recall asks what share of real events were flagged. Their denominators differ because they answer different operational questions. No one metric chooses a threshold for you: the costs of a miss, an unnecessary review, and the action triggered by a score belong in that decision too.

TypeScriptGo True positives · false positives · false negatives · true negatives · threshold tradeoffs
01 / Read the queue

A “suspicious” sign-in is a prediction, not a confirmed takeover.

For this lesson, the event is a sign-in attempt later confirmed to be an account takeover by the investigation team. The validation set contains 10,000 labeled attempts: 100 takeovers and 9,900 legitimate sign-ins. That 1% prevalence is illustrative. It gives us a population and a denominator; it is not a claim about any real authentication service.

The detector assigns a risk score, and the policy sends attempts above a chosen threshold to review. That creates four outcomes. A caught takeover is a true positive. A legitimate sign-in sent to review is a false positive. A missed takeover is a false negative. A legitimate sign-in left alone is a true negative. “False” describes the prediction relative to the current label; it does not tell us whether the threshold or the label process was reasonable.

Case file / Sign-in risk review All three candidate settings use the same 10,000 labeled attempts.
Event
A takeover is later confirmed for this sign-in attempt.
Population
100 confirmed takeovers and 9,900 legitimate attempts.
Score action
Send an attempt above threshold for investigation.
Decision
Which threshold supports the response policy and available capacity?
Question before the metric: what is the event, how was it labeled, and which sign-ins belong in this evaluation window?
02 / Count four outcomes

The confusion matrix is a count table, before it is a score.

At the middle threshold, the detector flags 278 attempts. Eighty are labeled takeovers and 198 are labeled legitimate. Among the 9,722 attempts it does not flag, 20 are takeovers and 9,702 are legitimate. Each number describes an intersection of actual condition and model prediction; all four cells must add back to the evaluated cohort.

These are constructed counts so the arithmetic is visible. In a real evaluation, labels might arrive days later, some takeovers might never be reported, and “legitimate” may include unrecognized compromise. Record the labeling rule and observation window with the matrix.

Middle threshold · rows are later-confirmed condition · columns are prediction
Actual conditionFlagged for reviewNot flaggedActual total
Takeover80 true positives20 false negatives100
Legitimate sign-in198 false positives9,702 true negatives9,900
Predicted total2789,72210,000
01 / Actual takeovers80 + 20 = 100

Flagged plus missed events.

02 / Legitimate attempts198 + 9,702 = 9,900

Unnecessary reviews plus correctly unflagged attempts.

03 / Entire cohort100 + 9,900 = 10,000

Rows and columns must reconcile to the same population.

03 / Choose the denominator

Keep the question attached to the fraction.

Precision is TP / (TP + FP): of all attempts flagged, what fraction were labeled takeovers? Here that is 80 / (80 + 198) = 80 / 278 ≈ 28.8%. The denominator is the flagged queue. Its complement, FP / (TP + FP), is the share of flags that were false positives under these labels.

Recall, also called sensitivity or true-positive rate, is TP / (TP + FN): of the actual takeovers in the cohort, what fraction did the detector flag? Here it is 80 / (80 + 20) = 80 / 100 = 80%. The denominator is actual takeovers, not all flags. Its complement, FN / (TP + FN), is the miss rate.

Accuracy is (TP + TN) / N, or the share of all predictions that match the labels. Here it is (80 + 9,702) / 10,000 = 97.82%. But if a detector flags nothing, it gets all 9,900 legitimate attempts “right” and misses all 100 takeovers: 99% accuracy, 0% recall. When positives are rare, a large true-negative count can dominate accuracy.

01 / Precision80 ÷ 278 ≈ 28.8%

TP divided by all flagged attempts.

02 / Recall80 ÷ 100 = 80%

TP divided by all actual takeovers.

03 / Accuracy9,782 ÷ 10,000 = 97.82%

Correct predictions divided by all attempts.

04 / Move the threshold

More catches usually put more legitimate sign-ins in the queue.

A score threshold converts a ranking into an action. Lower it and more attempts are flagged; raise it and fewer are. That often increases recall and lowers precision at the lower threshold, but an actual precision-recall curve must be measured from scored labeled data. The three settings below are an illustrative validation snapshot, not a guarantee that every model or population follows the same neat pattern.

Interactive lab / Fixed labeled cohort

Compare threshold consequences

A middle setting · still requires a policy choice. The cohort stays fixed at 100 labeled takeovers and 9,900 legitimate sign-ins.

Precision28.8%80 / 278 flagged attempts were takeovers
Recall80.0%80 / 100 takeovers were caught
Review queue278flagged attempts of 10,000 total
Middle threshold · each cell changes with the selected setting
Actual conditionFlaggedNot flaggedTotal
Takeover80 true positives20 false negatives100
Legitimate198 false positives9702 true negatives9,900
Predicted total278972210,000

These are illustrative counts, not live measurements. In production, select thresholds on a validation set, then evaluate on a separate time period or cohort before changing policy.

Same 10,000 attempts · illustrative counts at three score cutoffs
SettingTP / FPFNPrecisionRecallQueue size
Broad / lower94 / 89069.6%94%984
Middle80 / 1982028.8%80%278
Strict / higher62 / 353863.9%62%97
05 / Check the evidence

Before changing a cutoff, find out what moved.

A week-over-week precision drop could mean the score separates events less well. It could also mean takeovers are less prevalent, the score threshold changed, the analyst team labels more cases as legitimate, or a new region contributes different traffic. The metric is an observation; each explanation makes a different prediction. Check those predictions before turning the threshold knob.

Start with the raw four counts by week, then slice by region, device class, new versus returning account, and score band where sample sizes support a comparison. Verify that every attempt had a chance to be labeled: only investigating high-score attempts creates a selective-label problem, because missed takeovers in the unreviewed group remain unknown. Also record label delay; recent attempts may not have had enough time to become confirmed.

Competing explanations · look for the measurement that can separate them
Possible explanationWhat it predictsEvidence to inspect
Takeovers became rarerPrecision may fall while conditional recall and false-positive rate remain similar.Prevalence from matured labels; compare the same definition and time window.
Score separation changedAt the same threshold, both missed takeovers and legitimate flags may shift.Score distributions and confusion counts by label, model version, and cohort.
Label process changedRecent or unreviewed cases may be counted as legitimate before evidence matures.Label source, review coverage, appeal outcomes, and time to confirmation.
Traffic mix shiftedAggregate metrics move while within-region or device metrics stay steadier.Prevalence and rates by relevant cohort, with counts and uncertainty.
Diagnostic sequence: verify event and labels → reconcile counts → compare prevalence and threshold → inspect slices → choose the next measurement.
06 / Practice in code

Compute from counts and make empty denominators visible.

The implementations accept observed whole-number counts, validate that none are negative, and calculate precision, recall, and accuracy. For the middle setting they return precision about 0.288, recall 0.8, and accuracy 0.9782. The fractions are unitless; report them with the counts and evaluation population so they do not float free of the data they summarize.

If no attempt was flagged, precision has a zero denominator and is undefined—not zero. If there were no positive labels, recall is undefined. The TypeScript result uses null for those cases; Go uses explicit HasPrecision, HasRecall, and HasAccuracy flags. Neither example checks whether the labels are correct or the cohort represents deployment.

Compare the four-count calculation in TypeScript and Go.

Both implementations preserve undefined metrics when a denominator is zero.

TypeScriptPrecision, recall, and accuracy from a confusion matrix
metrics.ts
export type ConfusionMatrix = {
	truePositive: number;
	falsePositive: number;
	falseNegative: number;
	trueNegative: number;
};

export type ClassificationMetrics = ConfusionMatrix & {
	precision: number | null;
	recall: number | null;
	accuracy: number | null;
};

/** Calculate classification metrics from observed, non-negative whole-number counts. */
export function classificationMetrics(matrix: ConfusionMatrix): ClassificationMetrics {
	for (const [name, value] of Object.entries(matrix)) {
		if (!Number.isSafeInteger(value) || value < 0) {
			throw new Error(`${name} must be a non-negative safe integer`);
		}
	}

	const { truePositive: tp, falsePositive: fp, falseNegative: fn, trueNegative: tn } = matrix;
	const predictedPositive = tp + fp;
	const actualPositive = tp + fn;
	const total = tp + fp + fn + tn;
	if (!Number.isSafeInteger(total)) {
		throw new Error('sum of counts must be a safe integer');
	}
	return {
		...matrix,
		precision: predictedPositive === 0 ? null : tp / predictedPositive,
		recall: actualPositive === 0 ? null : tp / actualPositive,
		accuracy: total === 0 ? null : (tp + tn) / total
	};
}
GoPrecision, recall, and accuracy from a confusion matrix
metrics.go
package mathpractice

import (
	"errors"
)

type ConfusionMatrix struct {
	TruePositive  int64
	FalsePositive int64
	FalseNegative int64
	TrueNegative  int64
}

type ClassificationMetrics struct {
	ConfusionMatrix
	Precision    float64
	Recall       float64
	Accuracy     float64
	HasPrecision bool
	HasRecall    bool
	HasAccuracy  bool
}

// ClassificationMetricsFromCounts calculates ratios from observed counts.
// A Has* flag is false when the corresponding denominator is zero.
func ClassificationMetricsFromCounts(m ConfusionMatrix) (ClassificationMetrics, error) {
	counts := []int64{m.TruePositive, m.FalsePositive, m.FalseNegative, m.TrueNegative}
	for _, count := range counts {
		if count < 0 || count > (1<<53)-1 {
			return ClassificationMetrics{}, errors.New("counts must be between 0 and 2^53-1")
		}
	}
	// Each input is at most 2^53-1, so summing four remains within int64.
	predictedPositive := m.TruePositive + m.FalsePositive
	actualPositive := m.TruePositive + m.FalseNegative
	total := predictedPositive + m.FalseNegative + m.TrueNegative
	result := ClassificationMetrics{ConfusionMatrix: m}
	if predictedPositive > 0 {
		result.Precision = float64(m.TruePositive) / float64(predictedPositive)
		result.HasPrecision = true
	}
	if actualPositive > 0 {
		result.Recall = float64(m.TruePositive) / float64(actualPositive)
		result.HasRecall = true
	}
	if total > 0 {
		result.Accuracy = float64(m.TruePositive+m.TrueNegative) / float64(total)
		result.HasAccuracy = true
	}
	return result, nil
}
07 / Change the population

The same catch and false-alarm rates can produce different precision.

Keep the middle setting's recall at 80% and false-positive rate at 2%: 198 / 9,900 = 2%. Now evaluate 10,000 attempts where takeovers are only 0.1% of the population: 10 takeovers and 9,990 legitimate attempts. The model would be expected to flag 8 takeovers and miss 2, while also flagging about 9,990 × 0.02 = 199.8 legitimate attempts. These are expected counts, so a fraction is appropriate in the calculation; actual observed counts must be whole.

Expected precision becomes 8 / (8 + 199.8) ≈ 3.85%, even though recall and false-positive rate were held constant. The detector's precision is not a permanent property that transfers everywhere. It depends on the population mix as well as the measured conditional rates. In smaller subgroups, sampling variation can make the observed percentage jump around too.

Before transporting a scorecard to a new region, account age, device class, or season, ask for mature labels and the four counts within that target population. If prevalence changes, model the precision consequence; if recall or false-positive rate changes, investigate score or process drift. Do not let a new aggregate conceal a harmed subgroup.

01 / Events10,000 × 0.001 = 10

Takeovers in the transfer population.

02 / True positives10 × 0.80 = 8

Expected catches at 80% recall.

03 / False positives9,990 × 0.02 = 199.8

Expected flags among legitimate attempts.

04 / Precision8 ÷ 207.8 ≈ 3.85%

Expected true alerts divided by all expected flags.

Take the idea with you: report the event definition, threshold, four counts, metric denominators, cohort, time window, and what action a prediction triggers.