The formats zoo
FP8 E4M3 against E5M2, BF16 against FP16: range traded for precision. Then MX: 32 small numbers sharing one power-of-two scale.
Loading the animation…
Concept
With a fixed number of bits, every format makes the same trade: exponent bits buy range (how many binades), mantissa bits buy precision (how finely each binade is divided). The animation sends one value, , from far below 1 to far above it, and asks each format to store it.
- In the solid part of each bar (the normal range) the relative error stays put at a level set by the mantissa: about 0.2% for BF16, 3.1% for FP8 E4M3 and 6.2% for E5M2 at this value.
- In the hatched part (the subnormals) precision drains away one bit per binade, and below it the value is flushed to zero.
- Above the bar the value overflows: to infinity for IEEE formats, to NaN for E4M3 (non-saturating), and FP4, which has neither, saturates at its largest value.
FP8 E4M3 against E5M2. One more exponent bit doubles the number of binades (32 against 18) and costs one mantissa bit. The FP8 formats paper's recipe uses E4M3 for weights and activations, where precision matters and the range is known, and E5M2 for gradients, whose range is wide and unpredictable.
BF16 against FP16. BF16 is FP32 with the mantissa cut to 7 bits: with FP32's exponent it covers almost all of FP32's range, so a model trained in FP32 can switch to it without overflow or underflow, at 8 significant bits. FP16 has 11 significant bits but only 40 binades, so FP16 training scales the loss to keep small gradients from flushing to zero.
Loading the animation…
Concept
Four bits cannot cover the range of a whole tensor, but they can cover a block of neighbouring values if the block carries its own scale. That is the idea of the microscaling (MX) formats of the OCP MX specification: 32 elements of 8, 6 or 4 bits share one 8-bit scale, a pure power of two (E8M0), so an MXFP4 block costs 4.25 bits per value.
The animation runs the specification's conversion (section 6.3). Find the largest magnitude in the block; take the largest power of two not above it, and divide by the largest power of two the element type can hold (4 for FP4). That is the scale . Every element is stored as , rounded to the element type, and clamped if it lands beyond the type's largest value.
With an ordinary block, MXFP4 keeps all but 2 of the 32 values away from zero. Now pick the block with an outlier: one value of 24 drags the scale up to 4, every other element becomes a small fraction of the grid, and 19 of them fall into FP4's gap around zero. Outliers are the central problem of low-bit formats, and small blocks are one answer: an outlier can only spoil its own 32 neighbours. MXFP8 keeps all 32 (E4M3 has far more values near zero). Pick tiny values and compare the last readout: stored directly as FP4, all 32 would flush to zero; with the shared scale, the block keeps its shape.
Maths
An MX block represents with , . For element type with largest power of two (E4M3: ; E5M2: ; E3M2: ; E2M3 and E2M1: ; MXINT8: ), the conversion picks
so the largest scaled element lies in : at the top of the element's range, or just beyond its largest value (for E4M3, up to 512 against a maximum of 448), where it is clamped. An element much smaller than the block maximum, with below half the element type's smallest subnormal, is stored as zero; for FP4 that happens below somewhere between and of the block maximum, depending on where the maximum sits in its binade.
The dot product of two MX vectors factors the scales out, , so the hardware multiplies narrow elements and applies one scale per block.
Code
The conversion, cut from src/lib/num/model.ts (exponentOf is the exact of a double; mxElement rounds to nearest even and saturates):
const se =
amax === 0
? -127
: Math.max(-127, Math.min(127, exponentOf(amax) - f.emax_elem));
const X = pow2(se);
const t = v / X;
const u = mode === "sr" ? rng.u32() : 0;
const el = mxElement(t, mid, mode, u);