← Math in Practice
Concept Math behind AI

Derivatives, gradients, and the chain rule

Trace how a small parameter change moves a model's loss, then check what that local calculation can and cannot tell you.

A model review shows that a prediction missed its labeled target. One engineer asks which parameter mattered most; another wants to know whether changing it would actually improve the model. A derivative can answer a narrow, useful question: near the current parameter value, how quickly does this defined loss change as that parameter changes?

The judgment to keep

A derivative is local sensitivity. The chain rule carries sensitivity backward through composed calculations; a gradient collects those sensitivities for many parameters. These quantities describe a chosen model and loss at a chosen point. They do not, by themselves, explain a production failure or prove that a model change will help users.

TypeScriptGo Derivative · local sensitivity · gradient · chain rule · finite differences
01 / Read the model review

A prediction is wrong; the first task is to define what “wrong” means numerically.

Imagine a team inspecting a tiny scoring model during a training review. For one example, an input feature has value x = 2, the labeled target is y = 1, and the model predicts ŷ = wx + b. At the current settings w = 1.5 and b = 0, the score is 3. The team uses half squared error, L = ½(ŷ − y)², so this example's loss is 2.

These are deliberately small illustrative numbers, not benchmark data and not a complete training system. They let us trace the calculation all the way from a model parameter to a loss. That path is the teaching target: before asking which parameter to change, identify the quantity being changed and the exact output being measured.

Case file / One labeled exampleHow sensitive is this example's loss to the weight?
Input, target
x = 2, y = 1 (one toy training example).
Model score
ŷ = wx + b, currently 3.
Loss
L = ½(ŷ − y)² = 2.
Question
Near w = 1.5, how much does the loss change per unit change in w?
Separate the evidence: “loss is 2 on this example” is a calculation. “The model is generally poor” would require a representative evaluation set and a defined success criterion.
02 / Measure local sensitivity

The derivative is the slope at a point, not a promise about every possible change.

For a function f(w), the derivative at w is the limit of the change in output divided by the change in input as that input change approaches zero:

f′(w) = limh→0 [f(w + h) − f(w)] / h

For our fixed x, y, and b, the loss is L(w) = ½(wx + b − y)². At the current point, the score is 3 and the residual ŷ − y is 2. A small increase in w increases the score by x times that change: for a small weight change Δw, the score changes exactly by 2Δw. Near this point the loss changes approximately by 4Δw, or 4 loss units per weight unit.

01 / Predictionŷ = 1.5 × 2 + 0 = 3

Model score for the example.

02 / Residualŷ − y = 3 − 1 = 2

Prediction minus target.

03 / Sensitivity∂L/∂w = 2 × 2 = 4

Loss changes locally by about 4 per weight unit.

03 / Follow the chain

The chain rule multiplies the sensitivities along a composed path.

The weight does not enter the loss in one jump. It first affects the prediction, and the prediction affects the loss. Write that path as w → ŷ → L. The chain rule says dL/dw = (dL/dŷ)(dŷ/dw). In our model, dL/dŷ = ŷ − y = 2 and dŷ/dw = x = 2, so dL/dw = 2 × 2 = 4.

This factorization is useful because real models are compositions of many operations. Each operation contributes a local derivative; multiplying them follows how a small upstream change propagates to a downstream quantity. For many parameters and examples, software automatic differentiation applies these local rules systematically. The arithmetic still depends on the function and the point being evaluated.

01 / Parameterw = 1.5

Changing this weight changes the score.

02 / Scoreŷ = wx + b = 3

dŷ/dw = x = 2.

03 / LossL = ½(ŷ − y)² = 2

dL/dŷ = ŷ − y = 2.

04 / ChaindL/dw = 2 × 2 = 4

Multiply along the dependency path.

dL/dw = (dL/dŷ)(dŷ/dw) = (ŷ − y)x = (3 − 1) × 2 = 4

04 / Read the gradient

A gradient is a list of local sensitivities, one for each parameter.

The score depends on two parameters, w and b. We can differentiate the same loss with respect to each one. Since ŷ = wx + b, the score changes by x per unit of weight and by 1 per unit of bias. Applying the chain rule gives ∂L/∂w = (ŷ − y)x = 4 and ∂L/∂b = (ŷ − y) = 2.

Together these partial derivatives form the gradient ∇L = (∂L/∂w, ∂L/∂b) = (4, 2). It tells how the loss changes locally along each parameter axis. For a sufficiently small change (Δw, Δb), the first-order approximation is ΔL ≈ 4Δw + 2Δb. That is a local approximation; for large moves, curvature can make it inaccurate.

Partial derivatives at w = 1.5 and b = 0 for this one example
ParameterPath to lossDerivativeMeaning here
Weight ww → ŷ → L∂L/∂w = 4One small positive weight change raises this example's loss by about 4 times the change.
Bias bb → ŷ → L∂L/∂b = 2One small positive bias change raises this example's loss by about 2 times the change.
05 / Diagnose before tuning

A correct gradient answers a narrow question; it does not identify the production cause.

  1. Reproduce the observation. Preserve the example, model version, preprocessing, target, and reported metric.
  2. Check the objective. Confirm the loss formula, reduction (sum or mean), and which examples are included.
  3. Trace dependencies. Write the parameter-to-output path and verify each local derivative.
  4. Compare numerically. Perturb one parameter slightly and compare observed loss change with the local estimate.
  5. Evaluate transfer. Check representative held-out cases and product-facing outcomes before attributing improvement.
06 / Change the example

Compare the chain-rule result with a finite-difference estimate.

The lab keeps the model and loss fixed, then changes the input, target, parameters, or step size. It compares the analytic gradient with a central finite difference: evaluate the loss at a small positive and negative parameter offset, then divide their difference by twice the offset. With the default values, the weight derivative is exactly 4 analytically and approximately 4 numerically.

Current calculation
ŷ = wx + b = 3.0000ŷ − y = 2.0000L = ½(ŷ − y)² = 2.0000
∂L/∂wchain rule: 2.0000 × 2.0000 = 4.0000finite difference: 4.000000
∂L/∂bchain rule: 2.0000 × 1 = 2.0000finite difference: 2.000000

Model: ŷ = wx + b. Loss: L = ½(ŷ − y)². The central-difference estimate varies w or b alone while holding all other inputs fixed.

07 / Practice in code

Keep the forward calculation and its sensitivities visible.

Both snippets calculate one example's prediction, residual, half-squared loss, and analytic partial derivatives. Given x = 2, y = 1, w = 1.5, b = 0, they produce prediction 3, loss 2, and gradient (4, 2). They reject non-finite inputs and overflowed results so invalid arithmetic is visible. The code does not update parameters or implement a full training loop.

Compare the same calculation in TypeScript and Go.

Each computes the derivative of one half-squared-error example using the chain rule.

TypeScriptOne-example loss and gradient
derivatives.ts
/** One training example: prediction = weight * input + bias; loss = 1/2 * error^2. */
export function inspectExample(input: number, target: number, weight: number, bias: number) {
	if (![input, target, weight, bias].every(Number.isFinite)) {
		throw new RangeError('Inputs and parameters must be finite numbers.');
	}
	const prediction = weight * input + bias;
	const error = prediction - target;
	const loss = 0.5 * error ** 2;
	const dLossDWeight = error * input;
	const dLossDBias = error;
	if (![prediction, error, loss, dLossDWeight, dLossDBias].every(Number.isFinite)) {
		throw new RangeError('Calculation exceeded the finite number range.');
	}
	return { prediction, error, loss, dLossDWeight, dLossDBias };
}

const result = inspectExample(2, 1, 1.5, 0);
console.log(result); // prediction=3, error=2, loss=2, gradients=(4, 2)
GoOne-example loss and gradient
derivatives.go
package main

import (
	"errors"
	"fmt"
	"math"
)

type Example struct {
	Prediction float64
	Error      float64
	Loss       float64
	DWeight    float64
	DBias      float64
}

// InspectExample differentiates one half-squared-error example analytically.
func InspectExample(input, target, weight, bias float64) (Example, error) {
	values := []float64{input, target, weight, bias}
	for _, value := range values {
		if math.IsNaN(value) || math.IsInf(value, 0) {
			return Example{}, errors.New("inputs and parameters must be finite")
		}
	}
	prediction := weight*input + bias
	errValue := prediction - target
	loss := 0.5 * errValue * errValue
	dWeight := errValue * input
	dBias := errValue
	for _, value := range []float64{prediction, errValue, loss, dWeight, dBias} {
		if math.IsNaN(value) || math.IsInf(value, 0) {
			return Example{}, errors.New("calculation exceeded the finite number range")
		}
	}
	return Example{prediction, errValue, loss, dWeight, dBias}, nil
}

func main() {
	result, err := InspectExample(2, 1, 1.5, 0)
	if err != nil {
		panic(err)
	}
	fmt.Printf("prediction=%.1f error=%.1f loss=%.1f gradients=(%.1f, %.1f)\n",
		result.Prediction, result.Error, result.Loss, result.DWeight, result.DBias)
}
08 / Check the assumptions

The local slope depends on the model, the data, the objective, and the point.

We used one continuous linear score, one labeled example, and a smooth half-squared loss. Real models may be nonlinear, may combine many examples, and may use objectives with boundaries or nondifferentiable points. A reported derivative can also depend on preprocessing, regularization, sample weighting, and whether the loss is summed or averaged. Those details belong to the question being answered, not in footnotes after the number.

At a point where a function is not differentiable, a single ordinary derivative may not exist; software may use a defined subgradient convention for some operations. For a very large parameter move, the first-order estimate can be poor. And even a correctly computed gradient of a training loss is not evidence by itself that user-visible quality improves. That must be measured on relevant data and outcomes.

Transfer exercise

The team changes from one example to the mean loss over 100 examples.

What extra information do you need before predicting whether the mean-loss weight derivative is positive or negative? Name a finite-difference check you would run, and one validation measurement that would still be needed before claiming the user issue improved.