← Math in Practice
Concept Math behind AI

Normalization

A changed score may come from changed scale, not changed signal.

A relevance model starts ranking short support tickets differently after a training refresh. One feature is response time in milliseconds; another is a 0–1 flag. Before changing model weights, the team checks whether their scales and fitted statistics still mean what the model expects.

The judgment to keep

Normalization changes representation, not information. Identify the axis being scaled, fit statistics only on the training population, and reuse those saved statistics at serving time. A transformed value is interpretable only with its method and reference population.

TypeScriptGo Min-max scaling · z-scores · feature and vector scope · leakage · drift
01 / Investigate the score

A scale mismatch can hide inside a healthy looking pipeline.

The ranking model combines features such as response time, number of prior contacts, and a binary “contains attachment” flag. After a data pipeline change, response times arrive in seconds rather than milliseconds. That feature's values shrink by a factor of 1,000 while the model weights remain unchanged. The model still runs and returns scores, but those scores no longer reflect the scale it learned.

Use this deliberately small training sample for one feature: response time [10, 20, 30] ms. A new request takes 25 ms. We will calculate two common transforms using only these three training observations.

Case file / Ranking regressionDid the model change, or did the feature's units change?
Training values
10, 20, 30 ms
New observation
25 ms
Suspected change
One pipeline may now report seconds, or use freshly computed statistics.
Evidence to collect
Feature units, saved scaler version, training range, and serving transform output.
Bring to the review: raw values with units, training-fit parameters, transformed values, and model/scaler version identifiers.
02 / Choose what to scale

Min-max and z-scores put values into different reference frames.

Min-max scaling maps a training minimum to 0 and maximum to 1: (x − min) / (max − min). For 10, 20, 30, the 25 ms observation maps to (25−10)/(30−10)=0.75. It preserves ordering and relative spacing within the fitted range, but a future value can be below 0 or above 1. Clamping hides that out-of-range evidence, so do so only when the product meaning calls for it.

A z-score subtracts the training mean and divides by standard deviation: (x − μ) / σ. The mean is 20; using population standard deviation for this small worked set gives √(200/3) ≈ 8.165, so 25 ms maps to about 0.612 standard deviations above the mean. A z-score is not a probability and does not guarantee a normal distribution.

Min-max0.75

Position within the fitted minimum-to-maximum span.

Z-score+0.612

Distance from fitted mean in standard-deviation units.

Different statisticSame observation

Choose based on model contract and diagnostic goal.

Outside rangeDo not hide it

Monitor out-of-range values and drift.

03 / Fit the right population

“Normalize each feature” and “normalize each vector” are different operations.

Per-feature scaling calculates one set of statistics for each column across the training examples. Response time has its own mean and spread; retry count has another. This is common in tabular model pipelines. The learned transformation can be applied later to each new row.

Per-vector normalization instead rescales one whole vector at a time, often to unit length. That can be useful for comparing embedding direction with cosine similarity, or for making a row's total magnitude irrelevant. It also removes magnitude information: two vectors pointing the same way become equivalent even if one originally represented a much stronger signal.

Neural networks also use layer normalization or batch normalization, which normalize activations according to specific axes and runtime/training rules. Those are architectural operations with learned parameters and behavior defined by the model implementation; they are not interchangeable with a dataset's min-max scaler.

Dataset fitColumn by column

Reuse training statistics for each future record.

Vector normRow by row

May preserve direction while discarding magnitude.

Layer normInside a model

Axes and learned parameters are architectural choices.

Question firstWhich axis?

Document what observations and coordinates share statistics.

04 / Check the data path

Find where the reference population entered the pipeline.

A common evaluation leak occurs when a scaler is fitted on the full dataset before train/test splitting. Test-set values then influence the training transform. The remedy is to split first, fit on training data only, then apply the fixed transform to validation, test, and serving data. In cross-validation, fit separately within each training fold.

Constant features have a zero min-max range and zero standard deviation, so the formulas divide by zero. Decide whether to drop the feature, map it to a documented constant, or use a library's defined behavior. Missing values, outliers, and production drift need explicit policies too. A min-max scaler is especially sensitive to extreme training values; a z-score can also be pulled by them.

05 / Try a new observation

Change the value while keeping the training frame fixed.

Interactive lab / Response-time featureTraining values are fixed at 10, 20, and 30 ms.

Fitted training frame: min 10, max 30, mean 20, population σ ≈ 8.165.

Transform: (x − 10) / 20

Output: 0.750

What does the flag tell you? The value is outside this illustrative training range; it does not by itself prove an error or tell you how to clamp it.

06 / Practice in code

Fit on training data, then transform an incoming value.

The snippets use population standard deviation for a compact example, reject empty or non-finite training data, and surface constant features instead of silently dividing by zero. Production pipelines must additionally define missing-value handling and serialization of the fitted parameters.

Go
normalize.go
package main

import (
	"errors"
	"math"
)

type Summary struct{ Min, Max, Mean, StandardDeviation float64 }

func Fit(values []float64) (Summary, error) {
	if len(values) == 0 {
		return Summary{}, errors.New("provide at least one training value")
	}
	min, max, sum := values[0], values[0], 0.0
	for _, value := range values {
		if math.IsNaN(value) || math.IsInf(value, 0) {
			return Summary{}, errors.New("values must be finite")
		}
		if value < min {
			min = value
		}
		if value > max {
			max = value
		}
		sum += value
	}
	mean := sum / float64(len(values))
	variance := 0.0
	for _, value := range values {
		variance += (value - mean) * (value - mean)
	}
	return Summary{min, max, mean, math.Sqrt(variance / float64(len(values)))}, nil
}

func MinMax(value float64, summary Summary) (float64, error) {
	span := summary.Max - summary.Min
	if span == 0 {
		return 0, errors.New("cannot min-max scale a constant feature")
	}
	return (value - summary.Min) / span, nil
}

func ZScore(value float64, summary Summary) (float64, error) {
	if summary.StandardDeviation == 0 {
		return 0, errors.New("cannot z-score a constant feature")
	}
	return (value - summary.Mean) / summary.StandardDeviation, nil
}
07 / Make the change safely

Ship a scaler as part of the model contract.

When a feature changes from milliseconds to seconds, update one explicit unit contract and test a known request end to end. When the population itself shifts, evaluate a newly fitted transform with the model as one versioned change. Keep raw-feature checks, transformed-feature checks, and outcome metrics side by side so a plausible-looking score cannot conceal a broken conversion.

Record the fit population, axes, method, variance convention, handling for constants and out-of-range values, and fitted parameters. In an AI model, also distinguish input-feature preprocessing from normalization layers inside the network.

Decision rule: keep the scaler and model version together; alert on unit and distribution changes; compare labeled outcomes before changing the transform.
Read the design in your languages.

These choices apply to the comparisons throughout this story.