/about
About this site
Numerics Explained is about the numbers inside a model: how a format spends its bits, how rounding and accumulation go wrong, and the scaling and quantisation tricks that let 8-, 6- and 4-bit numbers work. Each chapter is built around an animation. It is the fifth of a family of companion sites, with the Transformer Decoder Explainer, LLM Inference Explained, LLM Architectures Explained, GPU Kernels Explained and Systolic Arrays Explained. The chapters link the matching slides of the Local LLM hosting and Google TPU series.
The numerics library
reference/numerics.py is written in plain Python, so that every step is one IEEE double operation, and src/lib/num/model.ts repeats it operation for operation. It implements:
- encoding and decoding for FP32, FP16, BF16, FP8 E4M3 and E5M2, FP6 E3M2 and E2M3 and FP4 E2M1, in every rounding mode IEEE 754-2019 defines plus stochastic rounding, with IEEE overflow or the OCP saturating mode (see the formats);
- the MX block formats (a shared E8M0 power-of-two scale per 32 elements, by section 6.3 of the MX specification) and NF4;
- integer quantisation, symmetric (absmax) and asymmetric (zero-point), per tensor, per channel and per group;
- summation in a format (naive, Kahan, pairwise), and GPTQ, AWQ-style scaling and SmoothQuant on small layers;
- LLM.int8()'s mixed-precision decomposition, NF4's derivation from normal quantiles, and one dot product in FP32, FP16, INT8 and MXFP4 with the energy of each operation.
The tiny transformer
Chapters 8 and 9 quantise a real network: the Transformer Decoder Explainer's tiny model (two blocks, 16 wide, a 64-character vocabulary), whose TypeScript is vendored unchanged into this site, with a plain-Python port in reference/tiny.py. It runs in your browser twice, as is and with its weights or its KV cache quantised, and the pages compare the two over four prompts. Its weights are random, rescaled from the explainer's so that its blocks matter; its text is gibberish, its arithmetic real. A test proves the hooked forward pass equals the explainer's exactly, and the Python and TypeScript runs agree to a relative 10⁻¹² (they share exp, log, sin, cos and tanh, which each language may round differently in the last place) and exactly in every top prediction.
How it is checked
- Against numpy and ml_dtypes. tests/python/test_numerics.py decodes every code of every format up to 16 bits and compares it with numpy (FP16) and ml_dtypes (BF16, FP8, FP6, FP4); encodes tens of thousands of values, ties and overflows included, and compares the codes; and checks the integer quantisers, the three summations and GPTQ against numpy implementations.
- Against an exact oracle. For every rounding mode and saturation mode, the encoder is compared with an oracle that lists the format's values and decides with exact rational arithmetic.
- Against the specifications. The tables of the OCP FP8 and MX specifications (largest, smallest normal and subnormal values, NaN and infinity codes, the conversion cases of FP8's Table 3), worked MX conversions, and bitsandbytes' NF4 decision thresholds.
- Exact parity. scripts/make_fixtures.py writes the reference's results, and the unit tests require the TypeScript port to reproduce all of them exactly: every code of every format up to 16 bits, and every encode case in all six modes, through SHA-256 digests of the bit patterns. CI fails if the fixtures are out of date.
- Animations from the model. Every animation draws a sequence of states the library computes; a frame is a pure function of one state. The tests set chosen frames of every animation and require the caption to match the caption built from the Python reference's state for that frame.
- Numbers in the prose are printed from the library when the page is built, not typed.
What is illustrative
- The data sets (the weight matrix, the activations, the MX blocks, the vectors summed) are generated from a seeded integer random number generator; the "normal" values are sums of 12 uniforms. They show the mechanisms, not a particular model's numbers.
- Bits per weight count FP16 scales (and, for zero-point, a zero point of the weight's width); real formats store their scales in several ways.
- The FP16 and BF16 sums round after every addition, as a scalar loop does; GPU kernels usually accumulate in FP32 and in their own order.
- The layers of chapters 6 to 8 are generated; the tiny transformer's weights are random; the KV-cache layouts are simplified (one scale over the whole prompt for per-channel keys); the energy figures are for a 45 nm process.
The animations
Every animation has play and pause, step back and forward, a scrub bar, speeds from 0.25× to 4× and reset; with the animation focused, Space plays or pauses and the arrow keys step. Each step has a one-line caption, also announced to screen readers. With reduce motion set in your system, nothing plays by itself. Animations pause when scrolled out of view. Colours come from Okabe and Ito's colour-blind-safe palette, the same in light and dark mode: one colour per bit field (sign purple, exponent sky blue, mantissa green, shared scale orange) and one per method (nearest-even blue, stochastic orange); an error is a warning colour.
Source
The code, the library and the tests are on GitHub (MIT licence). The design system is copied from the companion sites; the README records where each piece came from.