Quantisation basics
Integers and a scale: absmax and zero-point, and why one scale per group of weights beats one per tensor, on a weight heat map.
Loading the animation…
Concept
A floating-point format spends some of its bits on an exponent so that every value carries its own scale. Integer quantisation takes the opposite route: store each weight as a small integer, to for INT4, and keep one scale for a whole group of weights. The integer grid is uniform, so it wastes nothing on range it does not need, but its step is set by the largest value in the group.
The simplest choice is absmax: the scale is the group's largest magnitude divided by , so the largest weight lands exactly on the top code, zero is exactly representable, and everything else is rounded to the nearest step. The question is how big a group shares a scale.
The animation quantises one weight matrix (16 output channels by 64 inputs, illustrative) to INT4 five times:
- Per tensor, one scale for everything: the rows with small weights get a step sized for the largest weight in the matrix, and most of their weights round to a handful of codes. SQNR (signal to quantisation noise) is 9.1 dB.
- Per channel, one scale per output row: each row gets a step that fits it. 14.9 dB.
- Per group of 32, 16 and 8 weights along each row: a few input columns in this matrix are larger in every row, and a group that contains one pays for it; smaller groups confine the damage. 17.0 dB, 19.7 dB and 21.8 dB.
The scales are not free. Stored in FP16, they add 16 bits per group: 4.25 bits per weight per channel, 4.5 with groups of 32, 6 with groups of 8. LLM weight quantisers commonly settle on groups of 64 or 128, where the overhead is an eighth or a quarter of a bit. Watch the error map fade as the groups shrink, and try INT3, where every bit of SQNR is precious.
Loading the animation…
Concept
Absmax is symmetric about zero. Weights usually are too, but many activations are not: after a GELU or a ReLU they are mostly positive, and a symmetric grid then spends half its codes on negative values that hardly occur. Zero-point (asymmetric, or affine) quantisation shifts the grid to cover the values' actual range, from the minimum to the maximum, using all codes; the integer that represents 0.0 is the zero point.
In the animation, 24 skewed activations are quantised to INT4 both ways. The absmax grid has a step of 0.1952 and runs from to steps; the values use 7 of its 15 codes; the zero-point grid (zero point 2) uses 10 of 16, with a finer step of 0.1091, and its mean squared error is 0.001313 against absmax's 0.003169.
The price is arithmetic: with a zero point, a dot product expands into the integer product plus correction terms that involve the sums of the codes. Kernels precompute those terms, but it is one reason weights are often quantised symmetrically and zero points kept for activations.
Maths
For uniform quantisation with step and a value that is not clamped, the rounding error is at most , and if it is spread evenly over its mean square is . With absmax, , so each extra bit halves and quarters the noise: about dB of SQNR per bit,
The second term is where granularity acts: a group whose maximum is far above its typical value has a small , a coarse step for its size, and a low SQNR. Smaller groups raise it. With scales of bits per group of weights, the cost is bits per weight, plus the zero point's bits for asymmetric quantisation.
Affine quantisation (Jacob et al.) represents a real value as with an integer, so that is exact: zero padding and ReLU outputs stay exactly zero.
Code
The scale and zero point of a group, cut from src/lib/num/model.ts:
const [lo, hi] = intRange(bits, false);
const scale = mx > mn ? (mx - mn) / hi : 1;
const zero = Math.min(hi, Math.max(lo, rneInt(-mn / scale)));
return { scale, zero, lo, hi };
and the quantiser itself:
export function quantInt(x: number, p: IntParams): number {
const q = rneInt(x / p.scale) + p.zero;
return Math.min(p.hi, Math.max(p.lo, q));
}