numerics-explained

/formats

The formats, and where they are defined

Each format is four numbers (exponent bits, mantissa bits, bias, and how it spends the all-ones exponent). Everything else in these tables is derived from them by the tested library and agrees with its source's own table. The library decodes every code of every format up to 16 bits identically to numpy and ml_dtypes; see about for the checks.

Floating-point formats

FormatBits (s·e·m)BiasLargestSmallest normalSmallest subnormalPrecisionBinadesSpecialsDefined in
FP321·8·231273.40282 × 10³⁸2⁻¹²⁶2⁻¹⁴⁹24 bits (ulp at 1 = 2⁻²³)277±∞ and NaN (all-ones exponent)IEEE Standard for Floating-Point Arithmetic, Table 3.5, binary32
FP161·5·101565,5042⁻¹⁴2⁻²⁴11 bits (ulp at 1 = 2⁻¹⁰)40±∞ and NaN (all-ones exponent)IEEE Standard for Floating-Point Arithmetic, Table 3.5, binary16
BF161·8·71273.38953 × 10³⁸2⁻¹²⁶2⁻¹³³8 bits (ulp at 1 = 2⁻⁷)261±∞ and NaN (all-ones exponent)Google Cloud, 1 sign, 8 exponent, 7 mantissa bits; FP32's exponent
FP8 E5M21·5·21557,3442⁻¹⁴2⁻¹⁶3 bits (ulp at 1 = 2⁻²)32±∞ and NaN (all-ones exponent)OCP 8-bit Floating Point Specification (OFP8), Tables 1 and 2
FP8 E4M31·4·374482⁻⁶2⁻⁹4 bits (ulp at 1 = 2⁻³)18NaN only (S.1111.111); no ∞OCP 8-bit Floating Point Specification (OFP8), Tables 1 and 2
FP6 E3M21·3·23282⁻²2⁻⁴3 bits (ulp at 1 = 2⁻²)9none: every code is a numberOCP Microscaling Formats (MX) Specification, Table 4
FP6 E2M31·2·317.52⁰2⁻³4 bits (ulp at 1 = 2⁻³)6none: every code is a numberOCP Microscaling Formats (MX) Specification, Table 4
FP4 E2M11·2·1162⁰2⁻¹2 bits (ulp at 1 = 2⁻¹)4none: every code is a numberOCP Microscaling Formats (MX) Specification, Table 5

Binades: the powers of two from the smallest subnormal to the largest value, inclusive. For E4M3 and E5M2 this gives 18 and 32, as the OCP FP8 specification's Table 2 states.

MX block formats

A block of 32 elements shares one E8M0 scale X = 2k (k from −127 to 127; code 255 is NaN). The concrete formats of the OCP MX specification 1.0 (Table 1):

FormatElementElement bitsBlockScaleBits per valueLargest element power of two
MXFP8 (E4M3)FP8 E4M3832E8M08.2528
MXFP8 (E5M2)FP8 E5M2832E8M08.25215
MXFP6 (E3M2)FP6 E3M2632E8M06.2524
MXFP6 (E2M3)FP6 E2M3632E8M06.2522
MXFP4 (E2M1)FP4 E2M1432E8M04.2522
MXINT8INT8, 6 fraction bits (±1 63/64)832E8M08.2520

NF4 (NormalFloat 4)

Not a floating-point format: a code book of 16 values derived from quantiles of the normal distribution, scaled per block of 64 by the block's absolute maximum (QLoRA). The values are bitsandbytes' table; a value is stored as the nearest code, ties to the lower one.

CodeValueUpper decision threshold
0-1-0.8480964004993439
1-0.6961928009986877-0.6106329262256622
2-0.5250730514526367-0.4599952697753906
3-0.39491748809814453-0.33967943489551544
4-0.28444138169288635-0.23460740596055984
5-0.18477343022823334-0.13791173323988914
6-0.09105003625154495-0.045525018125772476
700.03979014977812767
80.079580299556255340.1202552504837513
90.160930201411247250.2035212516784668
100.246112301945686340.2920137718319893
110.337915241718292240.3893125355243683
120.440709829330444340.5016634166240692
130.56261700391769410.6427869200706482
140.72295683622360230.8614784181118011
151–

Rounding modes

The five modes of IEEE 754-2019 (section 4.3) and stochastic rounding. Overflow follows section 7.4 (towards-zero modes give the largest finite value), or the OCP saturating mode when it is chosen; formats without infinity saturate or give NaN, as their specifications say.

ModeRounds
roundTiesToEvento the nearer neighbour; a tie goes to the one whose last mantissa bit is 0 (IEEE 754's default, the only mode OCP requires)
roundTiesToAwayto the nearer neighbour; a tie goes away from zero
roundTowardZeroalways towards zero (truncation)
roundTowardPositivealways towards +∞
roundTowardNegativealways towards −∞
stochasticup with probability equal to the distance above the lower neighbour, in ulps (not an IEEE mode; MX allows other modes)

Sources

8 formats, 6 MX formats, 6 rounding modes.