numerics-explained
← /learn · 01

Bits to numbers

Sign, exponent and mantissa: click the bits of a floating-point code and watch its value move, then zoom in and see the values crowd towards zero.

Loading the animation…

Concept

A floating-point number is three fields of bits. The sign says which side of zero. The exponent says which power of two the number lies between: 232^3 and 242^4, say. The mantissa says where between them, in 2m2^m equal steps. Click the bits in the animation: an exponent bit moves the value by a power of two, a mantissa bit nudges it within its power of two.

The stretch from one power of two to the next is called a binade. Every binade holds the same 2m2^m values, but each binade is twice as wide as the one below it, so the gap between neighbours doubles every time the exponent goes up. That is why the format is called floating point: the relative precision is roughly constant (about m+1m+1 significant bits wherever you are) while the absolute spacing grows with the number. Press play and the number line zooms towards zero, a power of two per step: the same number of values keeps fitting into half the width, so they crowd together.

FP8 E4M3 (4 exponent bits, 3 mantissa bits) has 126 positive values, from 2⁻⁹ up to 448. Near 1 its neighbours are 0.125 apart; near 400 they are 32 apart.

Subnormals. When the exponent field is all zeros, the exponent stops at its minimum and the implicit leading 1 becomes a 0. The values below the smallest normal number, 2⁻⁶ for E4M3, are then evenly spaced down to zero instead of crowding further: this is IEEE 754's gradual underflow. The last zoom steps show it.

Infinity and NaN. IEEE formats (FP32, FP16, BF16 and FP8 E5M2) keep the all-ones exponent for ±∞\pm\infty and NaN, "not a number". FP8 E4M3 gives that binade back to ordinary numbers and keeps only S.1111.111 for NaN. The OCP FP8 specification does this "to increase emax to 8 and thus to increase the dynamic range by one binade", so the largest value is 448 rather than 240. The MX specification's FP6 and FP4 reserve nothing: every one of their codes is a number.

BF16 keeps FP32's 8-bit exponent and cuts the mantissa to 7 bits: nearly the same range as FP32 (up to about 3.4×10383.4 \times 10^{38}) with only 8 significant bits. FP16 has 11 significant bits but stops at 65,504. Chapter 4 puts them side by side; the formats page lists every constant.

Maths

For a format with ee exponent bits, mm mantissa bits and bias bb, a code with sign ss, exponent field EE and mantissa field MM is worth

v={(−1)s 2E−b(1+M2m)0<E (normal)(−1)s 21−b M2mE=0 (subnormal or zero)v = \begin{cases} (-1)^s \, 2^{E-b} \left(1 + \frac{M}{2^m}\right) & 0 < E \text{ (normal)} \\ (-1)^s \, 2^{1-b} \, \frac{M}{2^m} & E = 0 \text{ (subnormal or zero)} \end{cases}

with emin⁡=1−be_{\min} = 1 - b. In the binade [2k,2k+1)[2^k, 2^{k+1}) the spacing is the ulp (unit in the last place)

ulp⁡(x)=2max⁡(k, emin⁡)−m,2k≤∣x∣<2k+1,\operatorname{ulp}(x) = 2^{\max(k,\, e_{\min}) - m}, \qquad 2^k \le |x| < 2^{k+1},

so the gap relative to the number lies between 2−m−12^{-m-1} and 2−m2^{-m}, apart from the subnormals. The largest value is 2emax⁡(2−2−m)2^{e_{\max}}(2 - 2^{-m}) for formats that keep a full top binade, and 28×1.75=4482^{8} \times 1.75 = 448 for E4M3, whose top code is NaN.

Code

Decoding, cut from the site's src/lib/num/model.ts (the Python reference repeats it line for line, and the tests compare every code with numpy and ml_dtypes):

if (f.kind === "fn" && E === top && M === pow2(f.m) - 1) return NaN;
let v: number;
if (E === 0) v = pow2(f.emin - f.m) * M;
else v = pow2(E - f.bias - f.m) * (pow2(f.m) + M);
return s ? -v : v;