A training curve tells you what happened to one objective, not why.
Imagine a team retraining a ranking model after adding a new numeric feature. The training loss drops for several updates, then swings above and below its previous values. A larger learning rate might explain overshooting. So might a feature with a much larger numeric scale, a different batch mix, or noisy gradients. The curve is an observation; “the rate is too high” is a hypothesis that needs a test.
We will use a deliberately tiny model with one parameter and one loss value. This removes batches, many parameters, and data preparation so the arithmetic is visible. It teaches the update rule, not how to train a production ranking system.
- Observation
- Loss is measured on the current training batches; its units follow the chosen loss function.
- Change
- A numeric feature was added and transformed in the data pipeline.
- Competing explanations
- Step size, feature scale, batch noise, or a changed objective/data population.
- Question
- What does the optimizer do on each step, and what evidence would isolate the cause?
A one-dimensional bowl lets us check every step by hand.
Let the model have one parameter θ, with illustrative loss L(θ) = (θ − 3)². The loss is zero at θ = 3 and grows as θ moves
away. This is a convex parabola with one global minimum, so it is intentionally friendlier
than most real training objectives. We start at θ₀ = 0, where the loss is (0 − 3)² = 9.
The derivative is dL/dθ = 2(θ − 3). At θ = 0, the gradient is −6. Its
negative sign means that a small move toward larger θ lowers this loss. The slope
magnitude 6 says the loss changes steeply there per unit of θ. The derivative is not
itself a recommended parameter value: it is local information about how the loss changes.
Illustrative objective; minimum loss is 0.
Starting loss is 9.
At zero, increasing θ initially lowers loss.
The known minimum makes the trace easy to audit.
Subtract the gradient scaled by the learning rate.
For a parameter vector θ, objective L, and positive learning
rate η, the basic update is:
θₜ₊₁ = θₜ − η ∇L(θₜ)
With a single parameter, θ₀ = 0, gradient −6, and η = 0.1, the update is θ₁ = 0 − 0.1 × (−6) = 0.6. The new loss
is (0.6 − 3)² = 5.76, down from 9. For this quadratic, the derivative at every
new position points back toward 3.
In a model with many parameters, the gradient has one component per parameter. One learning rate scales all components in basic gradient descent. The resulting step can still be very different across directions because the loss surface may be steep in one direction and flat in another.
| Quantity | Calculation | Result |
|---|---|---|
| Current parameter | θ₀ | 0 |
| Gradient | 2(0 − 3) | −6 |
| Update | 0 − 0.1 × (−6) | θ₁ = 0.6 |
| New loss | (0.6 − 3)² | 5.76 |
Step size changes how quickly the parameter moves and whether it settles.
On this exact quadratic, the update simplifies to θₜ₊₁ − 3 = (1 − 2η)(θₜ − 3). The error from the minimum is multiplied by 1 − 2η every step. For 0 < η < 0.5, it approaches from the
same side. For 0.5 < η < 1, the sign flips each step, so it crosses the
minimum while shrinking its distance. At η = 1, it bounces between 0 and 6;
above 1, the distance grows on this loss.
That exact threshold is a property of this chosen curve and update rule, not a universal learning-rate recipe. In a real model, curvature, parameter scaling, minibatch noise, momentum, adaptive optimizers, clipping, and schedules all change the observed behavior. “Lower” is not automatically “better”: a tiny rate may make progress too slowly for a useful training budget.
| η | Parameter trace | Observed pattern |
|---|---|---|
| 0.05 | 0 → 0.3 → 0.57 → 0.813 → 1.0313 | Steady but slow approach to 3; more updates are needed. |
| 0.8 | 0 → 4.8 → 1.92 → 3.648 → 2.6112 | Crosses the minimum; oscillation shrinks on this curve. |
| 1.1 | 0 → 6.6 → −1.32 → 8.184 → −3.2208 | Overshoots with growing distance; this run diverges. |
Change one control and read the parameter and loss at every step.
This lab calculates the same exact quadratic, L(θ) = (θ − 3)², and its
derivative 2(θ − 3). Choose a rate, number of updates, and start value. The
table is a deterministic model with exact gradients; it contains no training data, batch
noise, or generalization result.
Loss is lower at this step count. Final θ = 2.8600; loss = 0.0196. The known minimum is θ = 3.
| Step | θ | Gradient | Loss |
|---|---|---|---|
| 0 | 0.0000 | -6.0000 | 9.0000 |
| 1 | 4.8000 | 3.6000 | 3.2400 |
| 2 | 1.9200 | -2.1600 | 1.1664 |
| 3 | 3.6480 | 1.2960 | 0.4199 |
| 4 | 2.6112 | -0.7776 | 0.1512 |
| 5 | 3.2333 | 0.4666 | 0.0544 |
| 6 | 2.8600 | -0.2799 | 0.0196 |
Make the update and its consequences visible in the data structure.
The examples calculate the same one-parameter trace and keep the starting point, gradient, and loss with each row. The guards validate finite inputs and a bounded step count; they do not claim to be a complete machine-learning optimizer. A production optimizer must also define parameter storage, data batches, gradient computation, numerical precision, and checkpoint behavior.
Both examples use the same quadratic, starting point, and learning-rate sweep.
export type Step = { iteration: number; theta: number; gradient: number; loss: number };
// A deliberately small, deterministic loss surface: L(theta) = (theta - 3)^2.
export function gradientDescent(learningRate: number, steps: number, start = 0): Step[] {
if (!Number.isFinite(learningRate) || learningRate <= 0) {
throw new RangeError('learningRate must be finite and greater than zero');
}
if (!Number.isInteger(steps) || steps < 0 || steps > 100) {
throw new RangeError('steps must be an integer from 0 to 100');
}
if (!Number.isFinite(start)) throw new TypeError('start must be finite');
const trace: Step[] = [
{ iteration: 0, theta: start, gradient: 2 * (start - 3), loss: (start - 3) ** 2 }
];
for (let iteration = 1; iteration <= steps; iteration++) {
const previous = trace[iteration - 1].theta;
const gradientAtPrevious = 2 * (previous - 3);
const theta = previous - learningRate * gradientAtPrevious;
trace.push({
iteration,
theta,
gradient: 2 * (theta - 3),
loss: (theta - 3) ** 2
});
}
return trace;
}
// Example: inspect the parameter and loss at every update for three rates.
export function compareLearningRates() {
return [0.05, 0.8, 1.1].map((rate) => ({
learningRate: rate,
trace: gradientDescent(rate, 6)
}));
}
package main
import (
"errors"
"fmt"
"math"
)
type Step struct {
Iteration int
Theta float64
Gradient float64
Loss float64
}
// A deliberately small, deterministic loss surface: L(theta) = (theta - 3)^2.
func GradientDescent(learningRate float64, steps int, start float64) ([]Step, error) {
if math.IsNaN(learningRate) || math.IsInf(learningRate, 0) || learningRate <= 0 {
return nil, errors.New("learningRate must be finite and greater than zero")
}
if steps < 0 || steps > 100 {
return nil, errors.New("steps must be from 0 to 100")
}
if math.IsNaN(start) || math.IsInf(start, 0) {
return nil, errors.New("start must be finite")
}
gradient := 2 * (start - 3)
trace := []Step{{Iteration: 0, Theta: start, Gradient: gradient, Loss: (start - 3) * (start - 3)}}
for iteration := 1; iteration <= steps; iteration++ {
previous := trace[iteration-1].Theta
gradient = 2 * (previous - 3)
theta := previous - learningRate*gradient
trace = append(trace, Step{
Iteration: iteration,
Theta: theta,
Gradient: 2 * (theta - 3),
Loss: (theta - 3) * (theta - 3),
})
}
return trace, nil
}
func main() {
for _, rate := range []float64{0.05, 0.8, 1.1} {
trace, err := GradientDescent(rate, 6, 0)
if err != nil {
panic(err)
}
fmt.Printf("learning rate %.2f:\n", rate)
for _, step := range trace {
fmt.Printf(" %d: theta=%.4f loss=%.4f\n", step.Iteration, step.Theta, step.Loss)
}
}
}
A falling training loss answers a narrower question than “will this work?”
Training loss is the objective evaluated on training examples, often batches sampled from that set. If it falls, the optimizer is reducing that measured objective under the current pipeline. This can confirm that updates are being applied and can show whether the chosen objective is improving. It cannot show that the examples represent future traffic, that the labels are correct, or that a product outcome improved.
A validation metric evaluates a separate held-out set and offers evidence about performance on examples not used to compute those updates. It is still not a guarantee about production: the validation data may be stale, unrepresentative, or repeatedly used to tune decisions. Keep a final test set or later time window for a less repeatedly consulted check, and evaluate the product metric that actually matters, with slices that can reveal regressions.
If training loss falls while validation loss rises, one plausible explanation is overfitting, but first check that the two losses use compatible definitions and that preprocessing is consistent. If both stall, possible causes include a low rate, poor features, a gradient bug, an objective mismatch, or a plateau in the loss surface. The curves narrow the search; they do not uniquely identify the cause.
The bowl is simple; real loss surfaces and gradients are not.
The learning-rate behavior above follows from a smooth, one-dimensional, deterministic quadratic and exact gradients. Real objectives have many parameters and can be non-convex. A local minimum is lower than nearby points but may not be globally best. A saddle point can have a nearly zero gradient while some directions descend and others ascend. Flat regions, steep narrow directions, and poorly scaled features can make one global rate inefficient or unstable.
With minibatch training, the gradient is an estimate based on the selected examples, so different batches can point in slightly different directions. This noise may help exploration, but it also makes individual loss values fluctuate. Large gradient norms, non-finite values, clipping, regularization, and optimizer state can all affect the trace. A plotted curve may also average or smooth values, so inspect how it was aggregated.
For the incident, make a diagnostic sequence: reproduce the preprocessing ranges; compare gradient norms and update magnitudes before and after the feature change; hold initialization, batches, and objective fixed while sweeping the rate; then compare training and validation curves and representative slices. If the feature scale changed dramatically, a controlled normalization experiment can test that hypothesis. Change one factor at a time and keep the experiment record, rather than selecting the run that merely looks smooth.
References
- Google Machine Learning Crash Course: Gradient descent — learning-rate intuition and iterative parameter updates.
- Google Machine Learning Crash Course: Overfitting — training and validation behavior as distinct evidence.