numerics-explained

/learn

Number formats, rounding and quantisation

One chapter per idea, each opening with an animation. Every frame is computed by this site's numerics library (checked bit for bit against its Python reference, numpy and ml_dtypes), and every format follows its published specification. Toggle layers (Concept / Maths / Code) inside any chapter to choose how deep to go. How a GPU runs quantised kernels is on GPU Kernels Explained; for slides, see the Local LLM hosting and Google TPU series.

  1. Sign, exponent and mantissa: click the bits of a floating-point code and watch its value move, then zoom in and see the values crowd towards zero.

  2. Round to nearest, ties to even, against stochastic rounding: one is the most accurate on each value, the other the only one that is right on average.

  3. Summing 8,192 numbers in FP16 stalls at 2,048. FP32, Kahan's compensated sum and pairwise summation, raced live.

  4. FP8 E4M3 against E5M2, BF16 against FP16: range traded for precision. Then MX: 32 small numbers sharing one power-of-two scale.

  5. Integers and a scale: absmax and zero-point, and why one scale per group of weights beats one per tensor, on a weight heat map.

  6. A few activation channels far larger than the rest wreck per-tensor INT8. LLM.int8() splits them off; SmoothQuant moves the difficulty into the weights.

  7. Quantise one column, push its error into the columns still to come, weighted by the inverse Hessian: second-order error compensation, column by column.

  8. Scale the input channels that matter before rounding; build a 4-bit code book from the normal distribution's quantiles. Then quantise a real (tiny) transformer.

  9. Keys and values in 8, 4 and 2 bits, per token and per channel, and FP4 with a scale per 16: memory saved against logit drift, measured on a live transformer.

  10. Where the bits pay off: dequantising in registers, block-scaled formats in tensor cores, and the energy and area of each operation.