Numerics Explained
How few bits can a number have?
Models are trained in 16-bit formats and served in 8, 6 or 4. An FP8 E4M3 number has 4 significant bits and nothing above 448; FP4 has 2 and stops at 6. This site is about what that costs and the tricks that make it work: rounding, accumulating in wider formats, scales shared by blocks of numbers, and quantisers that choose their integers with care.
Each chapter is built around an animation, and every frame is computed by a small numerics library that reproduces numpy and ml_dtypes bit for bit and follows the IEEE 754 and OCP FP8 and MX specifications.
01
Bits to numbers
Sign, exponent and mantissa: click the bits of a floating-point code and watch its value move, then zoom in and see the values crowd towards zero.
02
Rounding
Round to nearest, ties to even, against stochastic rounding: one is the most accurate on each value, the other the only one that is right on average.
03
Accumulation error
Summing 8,192 numbers in FP16 stalls at 2,048. FP32, Kahan's compensated sum and pairwise summation, raced live.
04
The formats zoo
FP8 E4M3 against E5M2, BF16 against FP16: range traded for precision. Then MX: 32 small numbers sharing one power-of-two scale.
05
Quantisation basics
Integers and a scale: absmax and zero-point, and why one scale per group of weights beats one per tensor, on a weight heat map.
06
The outlier problem
A few activation channels far larger than the rest wreck per-tensor INT8. LLM.int8() splits them off; SmoothQuant moves the difficulty into the weights.
07
GPTQ
Quantise one column, push its error into the columns still to come, weighted by the inverse Hessian: second-order error compensation, column by column.
08
AWQ and NF4
Scale the input channels that matter before rounding; build a 4-bit code book from the normal distribution's quantiles. Then quantise a real (tiny) transformer.
09
Quantising the KV cache
Keys and values in 8, 4 and 2 bits, per token and per channel, and FP4 with a scale per 16: memory saved against logit drift, measured on a live transformer.
10
Quantisation in hardware
Where the bits pay off: dequantising in registers, block-scaled formats in tensor cores, and the energy and area of each operation.
Every format, its constants and the specification that defines it: the formats.
Part of a family of companion sites: the Transformer Decoder Explainer (one forward pass), LLM Inference Explained (serving it), LLM Architectures Explained (how the models differ), GPU Kernels Explained (how a GPU runs the maths) and Systolic Arrays Explained (the matrix hardware). This site is about the numbers themselves. How it was built, and how to check it: about.