The candidate with the largest score is easy to name; its probability takes another step.
A language model assigns next-token logits to the words or token fragments that could
follow. For this small example, three candidates have logits [2.4, 1.1, 0.2].
A logit is an unnormalized score: adding the same constant to every logit leaves their
relative ranking unchanged, and the values do not have to lie between zero and one or sum
to anything.
The question from the incident is empirical: did the model's scores change, did the temperature change, or did another decoding control alter token choice? Capture the logits and decoding configuration for the same prompt before changing settings. This lesson isolates the softmax stage so that each effect can be checked.
- Observed
- Generated text became more varied after a deployment.
- Candidate scores
- Three illustrative logits: 2.4, 1.1, and 0.2.
- Competing explanations
- Changed logits, temperature, top-k/top-p filtering, or random sampling.
- First check
- Compare captured logits and decoding configuration on the same prompt.
The exponential makes every weight positive and amplifies score differences.
For a logit zᵢ, exponentiation gives the unnormalized weight eᶻⁱ. Because exponentials are positive, the weights can be normalized. Their
ratio makes the role of differences visible: for two candidates, eᶻ¹ / eᶻ² = e⁽ᶻ¹⁻ᶻ²⁾. A difference of 1.3 means a weight ratio of e¹·³ ≈ 3.67, no matter whether the logits are 2.4 and 1.1 or both shifted upward by 100.
With the three example logits and temperature T = 1, subtract the maximum
first: [2.4, 1.1, 0.2] − 2.4 = [0, −1.3, −2.2]. The stable weights are [1, e⁻¹·³, e⁻²·²] ≈ [1, 0.2725, 0.1108]. The denominator is their sum,
approximately 1.3833.
Subtract it from each score without changing score differences.
Positive values preserve which candidate ranked higher.
Every weight will be divided by this shared total.
Dividing each weight by the total produces probabilities that sum to one.
At temperature T > 0, softmax is pᵢ = exp(zᵢ/T) / Σⱼ exp(zⱼ/T). For this example at T = 1, the
probabilities are approximately [0.723, 0.197, 0.080]. The first candidate
has the largest modeled share of next-token probability mass; the values sum to one apart
from display rounding.
Adding the same constant c to every logit does not change the result because exp((zᵢ+c)/T) contains a shared factor exp(c/T), which cancels
between numerator and denominator. This is useful diagnostically: large-looking absolute
logits alone do not tell us the distribution's sharpness; relative gaps and temperature
do.
| Candidate | Logit | Weight after max subtraction | Probability |
|---|---|---|---|
| A | 2.4 | exp(0) = 1 | 0.7229 |
| B | 1.1 | exp(−1.3) ≈ 0.2725 | 0.1970 |
| C | 0.2 | exp(−2.2) ≈ 0.1108 | 0.0801 |
| Total | — | 1.3833 | 1.0000* |
*Each displayed probability is rounded independently; the unrounded probabilities sum to one.
Temperature changes the spread of probability, not the ordering of finite logits.
Dividing by T scales each difference from the maximum. At T < 1, gaps grow before exponentiation, so the top candidate gets a larger
share. At T > 1, gaps shrink, so probability mass spreads toward the other
candidates. As positive T approaches zero, the distribution concentrates on the
maximum-logit candidate; as T grows very large, a finite candidate set approaches a
uniform distribution.
Temperature does not reorder candidates when logits are finite and T is positive. But it can change which token a random sampler emits. Top-k or nucleus (top-p) filtering can then remove candidates and renormalize a different set. To diagnose a generation change, compare these settings separately rather than attributing all variety to temperature alone.
Compute relative exponentials so a large shared score cannot overflow.
A direct calculation can try to evaluate exp(1000), which exceeds ordinary
floating-point range even though the final probabilities are well-defined. Let m = max(z). Then exp(zᵢ/T) / Σexp(zⱼ/T) = exp((zᵢ−m)/T) / Σexp((zⱼ−m)/T) because the same
factor exp(−m/T) cancels. Every exponent is now non-positive and the largest is zero,
so at least one weight is exactly one.
The corresponding log denominator uses log-sum-exp: log Σexp(zᵢ/T) = m/T + log Σexp((zᵢ−m)/T). This is the stable route for
log-probabilities and log-likelihoods too. The next lesson, Log-probabilities and information theory, uses this connection to explain surprisal, cross-entropy, and perplexity.
Move temperature while holding candidate scores fixed.
These controls change only the temperature. The logits remain [2.4, 1.1, 0.2], so any visible change comes from rescaling the same score gaps and applying softmax.
The bars represent next-token probability mass for this illustrative candidate set.
A temperature change alters the distribution, not the candidate scores or their rank. This lab has no random sampling step.
Keep score transformation and probability normalization explicit.
The examples validate finite logits and a positive finite temperature, subtract the maximum before exponentiating, and return probabilities. The TypeScript companion also shows a log-sum-exp helper. They do not implement token filtering, random sampling, or model inference.
Both examples use the same three logits and compare temperatures 0.5, 1, and 2.
export type Probability = { label: string; logit: number; probability: number };
/** Softmax for finite logits. Subtracting the maximum preserves probabilities and avoids overflow. */
export function softmax(logits: readonly number[], temperature = 1): Probability[] {
if (logits.length === 0) throw new RangeError('logits must not be empty');
if (!Number.isFinite(temperature) || temperature <= 0) {
throw new RangeError('temperature must be finite and greater than zero');
}
if (logits.some((logit) => !Number.isFinite(logit))) {
throw new TypeError('every logit must be finite');
}
const maximum = logits.reduce((current, logit) => Math.max(current, logit), -Infinity);
const weights = logits.map((logit) => Math.exp((logit - maximum) / temperature));
const total = weights.reduce((sum, weight) => sum + weight, 0);
return logits.map((logit, index) => ({
label: `Candidate ${index + 1}`,
logit,
probability: weights[index] / total
}));
}
/** log(sum(exp(logits / temperature))) using the same max-subtraction trick. */
export function logSumExp(logits: readonly number[], temperature = 1): number {
if (logits.length === 0) throw new RangeError('logits must not be empty');
if (!Number.isFinite(temperature) || temperature <= 0) {
throw new RangeError('temperature must be finite and greater than zero');
}
if (logits.some((logit) => !Number.isFinite(logit))) {
throw new TypeError('every logit must be finite');
}
const maximum = logits.reduce((current, logit) => Math.max(current, logit), -Infinity);
const scaledSum = logits.reduce(
(sum, logit) => sum + Math.exp((logit - maximum) / temperature),
0
);
return maximum / temperature + Math.log(scaledSum);
}
export const rankingLogits = [2.4, 1.1, 0.2] as const;
package main
import (
"errors"
"fmt"
"math"
)
type Probability struct {
Logit float64
Probability float64
}
// Softmax subtracts the maximum logit before exponentiating to avoid overflow.
func Softmax(logits []float64, temperature float64) ([]Probability, error) {
if len(logits) == 0 {
return nil, errors.New("logits must not be empty")
}
if math.IsNaN(temperature) || math.IsInf(temperature, 0) || temperature <= 0 {
return nil, errors.New("temperature must be finite and greater than zero")
}
maximum := logits[0]
for _, logit := range logits {
if math.IsNaN(logit) || math.IsInf(logit, 0) {
return nil, errors.New("every logit must be finite")
}
if logit > maximum {
maximum = logit
}
}
weights := make([]float64, len(logits))
total := 0.0
for i, logit := range logits {
weights[i] = math.Exp((logit - maximum) / temperature)
total += weights[i]
}
probabilities := make([]Probability, len(logits))
for i, logit := range logits {
probabilities[i] = Probability{Logit: logit, Probability: weights[i] / total}
}
return probabilities, nil
}
func main() {
logits := []float64{2.4, 1.1, 0.2}
for _, temperature := range []float64{0.5, 1, 2} {
probabilities, err := Softmax(logits, temperature)
if err != nil {
panic(err)
}
fmt.Printf("temperature %.1f:", temperature)
for _, item := range probabilities {
fmt.Printf(" %.4f", item.Probability)
}
fmt.Println()
}
}
A normalized distribution is internally consistent; it can still be wrong about the world.
Softmax always assigns probabilities over the candidate set it receives. That property does not show that the set contains the right answer, that model probabilities match real frequencies, or that the model's text is truthful. A model can be sharply concentrated on a wrong next token. Correctness is an outcome to measure; calibration asks whether predictions assigned a given probability are correct at approximately that rate over an appropriate population.
To investigate the repetitive-output incident, reproduce a fixed prompt set, record raw logits and all decoding controls, then change one control at a time. Compare token diversity and answer-level evaluation against reviewed outcomes. If logits themselves changed, inspect model version, prompt construction, tokenization, and input data. If only temperature changed, expect a changed sampling distribution while preserving score rank. Document filtering because it changes the candidate set seen by a sampler.
References
- Google Machine Learning Crash Course: Softmax function — maps logits to a normalized probability distribution.
- PyTorch: Softmax — definition and dimension-wise normalization.