Smaller weights change both the footprint and the computation.
A team is testing a model conversion to fit more concurrent requests into available device memory. The converted model is smaller, but a few outputs changed and one benchmark regressed. That does not yet identify whether the cause is overflow, rounding, a calibration range that clips activations, unsupported hardware behavior, or a task-specific sensitivity.
Begin with two separate questions. First, what storage reduction should the format provide? Second, what numerical and task changes does this exact model conversion cause on the target runtime? Parameter count answers the first roughly; evaluation and hardware measurements answer the second.
- Goal
- Reduce weight memory to fit a larger batch on target hardware.
- Risk
- Numerical error may concentrate in sensitive layers or outputs.
- Knowns
- Parameter count, storage format, calibration data, target runtime and device.
- Decision evidence
- Peak memory, latency/throughput, representative task metrics, and output/error slices.
fp16 has tighter range; bf16 keeps range with coarser steps.
IEEE-style floating-point values divide bits among sign, exponent, and significand. The exponent largely controls dynamic range; significand bits control how finely values are spaced near a magnitude. “16-bit” alone does not tell you the tradeoff.
fp32 uses 32 bits, about 24 bits of significand precision, and a maximum
finite value near 3.4×10³⁸. fp16 uses 16 bits and about 11
bits of significand precision; its maximum finite value is 65,504, so large intermediate
values can overflow even when fp32 handled them. bf16 also uses 16 bits,
with about 8 significand bits, but an exponent width similar to fp32 gives it a much wider
range, around 3.4×10³⁸. It represents a broad range coarsely.
int8 is a signed 8-bit integer with codes −128 through 127. It does not encode arbitrary real numbers by itself. A quantizer maps a chosen real interval to those codes using a scale and often a zero point. Its range depends on calibration and quantization scheme; values outside the interval may clip.
More significand precision for stored values.
Fine precision relative to bf16, lower max magnitude.
fp32-like exponent range, coarser significand.
Resolution depends on scale and represented interval.
Storage bytes are a useful first estimate, not peak serving memory.
For one million parameters, raw weight storage is about 4,000,000 bytes (4.00 MB or 3.81 MiB) in fp32, 2.00 MB in fp16 or bf16, and 1.00 MB in int8. This idealized arithmetic excludes metadata such as scales, zero points, alignment, and container overhead.
Runtime memory also includes activations, temporary workspaces, caches, allocator fragmentation, and possibly a higher-precision copy of some tensors. An int8 checkpoint does not guarantee an int8 execution path. Similarly, shorter arithmetic may increase throughput only when the target hardware and kernels support it efficiently.
4 bytes per value
2 bytes per value
1 byte per value
Runtime overhead is not included.
Quantization maps a real interval to discrete codes.
For a simple uniform signed int8 mapping of a calibration interval [low, high], the step size is approximately (high−low)/255. Each value is rounded to a
code; reconstruction maps the code back to a nearby real number. Widen the range and each
step gets coarser; narrow it and more outliers may clip. Real schemes can be symmetric or
asymmetric, per-tensor or per-channel, and may treat zero points and endpoints
differently.
Calibration data estimates ranges or scales. It should represent expected serving inputs, but it must remain separate from the held-out evaluation set used to decide whether the converted model meets its task goal. Evaluate both common and important rare slices: average error or an aggregate score can hide a failure in a sensitive category.
Values outside it may saturate or clip.
Wider interval at fixed bits means larger steps.
Mismatch can cause clipping or waste precision.
Check outputs on representative labeled data.
See clipping and reconstruction error separately.
Input values: -1.2, -0.65, -0.1, 0.1, 0.49, 1, 1.2
Step ≈ 0.0078; clipped values 2; reconstruction RMSE 0.1069.
| Original | int8 code | Reconstructed |
|---|---|---|
| -1.20 | -128 | -1.000 |
| -0.65 | -83 | -0.647 |
| -0.10 | -13 | -0.098 |
| 0.10 | 12 | 0.098 |
| 0.49 | 62 | 0.490 |
| 1.00 | 127 | 1.000 |
| 1.20 | 127 | 1.000 |
This is a transparent toy affine mapping. It does not simulate fp16/bf16 rounding, a neural network, or any particular framework's quantizer.
Pair numerical checks with the product task and serving target.
Keep a full-precision baseline, fix evaluation prompts/examples and decoding settings, then compare the converted model. Depending on the task, inspect exact-match or ranking quality, calibration, safety, structured output validity, or human-rated quality. Also inspect output deltas and error by layer or input slice where available. A small average error in weights or activations does not guarantee a small change in model behavior.
Benchmark the actual target device and runtime. Measure peak memory, throughput at a stated batch/concurrency, and latency percentiles after warmup. Include load/unload and conversion costs if they affect the service. Check for fallback operations that silently run in a wider format or on a different execution path.
Make the chosen interval visible in the calculation.
These snippets implement one educational uniform signed int8 mapping and a raw parameter-storage estimate. They validate inputs and report clipping and RMSE. They are not an implementation of floating-point formats or a substitute for your model framework's quantization workflow.
package main
import (
"errors"
"math"
)
type Quantized struct {
Values []float64
Scale float64
Clipped int
RMSE float64
}
func QuantizeInt8(values []float64, low, high float64) (Quantized, error) {
if len(values) == 0 || math.IsNaN(low) || math.IsInf(low, 0) || math.IsNaN(high) || math.IsInf(high, 0) || low >= high {
return Quantized{}, errors.New("use finite values and calibration range low < high")
}
for _, value := range values {
if math.IsNaN(value) || math.IsInf(value, 0) {
return Quantized{}, errors.New("values must be finite")
}
}
scale := (high - low) / 255
restored := make([]float64, len(values))
clipped := 0
squaredError := 0.0
for i, value := range values {
bounded := math.Max(low, math.Min(high, value))
if bounded != value {
clipped++
}
code := math.Round((bounded-low)/scale) - 128
reconstructed := (code+128)*scale + low
restored[i] = reconstructed
delta := value - reconstructed
squaredError += delta * delta
}
return Quantized{restored, scale, clipped, math.Sqrt(squaredError / float64(len(values)))}, nil
}
func StorageBytes(parameters, bytesPerValue int64) (int64, error) {
if parameters < 0 || bytesPerValue < 1 {
return 0, errors.New("use nonnegative parameter count and positive byte width")
}
return parameters * bytesPerValue, nil
}
These choices apply to the comparisons throughout this story.