numerics-explained

Numerics Explained

How few bits can a number have?

Models are trained in 16-bit formats and served in 8, 6 or 4. An FP8 E4M3 number has 4 significant bits and nothing above 448; FP4 has 2 and stops at 6. This site is about what that costs and the tricks that make it work: rounding, accumulating in wider formats, scales shared by blocks of numbers, and quantisers that choose their integers with care.

Each chapter is built around an animation, and every frame is computed by a small numerics library that reproduces numpy and ml_dtypes bit for bit and follows the IEEE 754 and OCP FP8 and MX specifications.

FP4 E2M1: 8 values, 0 to 6FP6 E2M3: 32 values, 0 to 7.5FP8 E4M3: 127 values, 0 to 448
Every value of three small formats, to scale: the gaps double at each power of two.

Every format, its constants and the specification that defines it: the formats.

Part of a family of companion sites: the Transformer Decoder Explainer (one forward pass), LLM Inference Explained (serving it), LLM Architectures Explained (how the models differ), GPU Kernels Explained (how a GPU runs the maths) and Systolic Arrays Explained (the matrix hardware). This site is about the numbers themselves. How it was built, and how to check it: about.