The outlier problem
A few activation channels far larger than the rest wreck per-tensor INT8. LLM.int8() splits them off; SmoothQuant moves the difficulty into the weights.
Loading the animation…
Concept
Chapter 5 sized a scale for the largest value in its group. That works while the largest value is typical. In the activations of large transformers it often is not: a handful of hidden dimensions carry values far larger than the rest, in almost every token. Dettmers et al. report that they emerge in large language models at around 6.7 billion parameters, systematic (the same few dimensions, token after token) and sparse.
The animation builds such an input: 16 channels × 32 tokens, where 2 channels reach 24.63 while no other channel exceeds 2.922. Then it multiplies the input by a small weight matrix three ways:
- Per-tensor INT8: one scale for the whole input, set by the outliers. The ordinary channels are left with 31 of the 255 codes, under five bits' worth. Output error 7.09 × 10⁻⁴.
- Vector-wise INT8, LLM.int8()'s first idea: one scale per token (row of the input) and one per output channel. It does not help here, because every token contains the outliers: 6.79 × 10⁻⁴.
- Mixed-precision decomposition: find the channels with any value above a threshold (LLM.int8() uses 6), multiply those columns in 16-bit and everything else in vector-wise INT8, and add. Error 3.46 × 10⁻⁶, 196.4 times smaller, while only 2 of 16 channels run in 16-bit. In real models the outlier dimensions are a fraction of a percent of the total.
The price is a second, irregular matrix multiply and the bookkeeping to split and merge. The paper's aim is memory: it reports that the overhead slows inference below 6.7B parameters, and that for matrix multiplications the size of a 175B model's it runs about twice as fast as FP16.
Loading the animation…
Concept
SmoothQuant (Xiao et al.) keeps everything in INT8 by changing the problem instead. Activations are hard to quantise because their channels differ wildly; weights are easy because theirs do not. Dividing activation channel by a factor and multiplying row of the weights by the same leaves the product exactly as it was, in real arithmetic, and the division can be folded into the previous layer's weights, so it costs nothing at run time.
The migration strength says how much of the difficulty to move. At the activations keep it all; at the weights take it all; at , the paper's default for most models, each channel's largest activation and largest weight become equal. The animation sweeps in eighths over the same layer. Per-tensor W8A8 error falls from 7.09 × 10⁻⁴ with no smoothing to 1.6 × 10⁻⁴ at , 4.441 times smaller; on this layer the best eighth is 0.625, with 1.46 × 10⁻⁴. Too far, and the weights' columns become the outliers.
Maths
With per-tensor absmax INT8, a channel whose values span uses a fraction of the code range. Its rounding noise is set by the global step , so its signal-to-noise ratio falls by dB against a channel that owns the whole range. A channel twenty times smaller than the outlier loses dB, more than four bits at about 6 dB per bit.
SmoothQuant's scaling makes the activation channel maxima and the weight row maxima , where . At both equal : the spread of the activation maxima is replaced by the square root of the spread of .
Code
LLM.int8()'s decomposition, cut from src/lib/num/model.ts:
const outl: number[] = [];
for (let i = 0; i < d; i++)
for (let k = 0; k < n; k++)
if (Math.abs(X[i]![k]!) > threshold) {
outl.push(i);
break;
}
const normal = Array.from({ length: d }, (_, i) => i).filter(
(i) => !outl.includes(i),
);
const Y8 = int8Vectorwise(W, X, normal);
and SmoothQuant's scales, with built from square roots so the Python reference and this port agree bit for bit:
const s = Array.from(
{ length: d },
(_, j) => powDyadic(xmax[j]!, k8) / powDyadic(wmax[j]!, 8 - k8),
);
const Xs = X.map((row, j) => row.map((x) => x / s[j]!));
const Ws = W.map((row) => row.map((w, j) => w * s[j]!));