/formats
The formats, and where they are defined
Each format is four numbers (exponent bits, mantissa bits, bias, and how it spends the all-ones exponent). Everything else in these tables is derived from them by the tested library and agrees with its source's own table. The library decodes every code of every format up to 16 bits identically to numpy and ml_dtypes; see about for the checks.
Floating-point formats
| Format | Bits (s·e·m) | Bias | Largest | Smallest normal | Smallest subnormal | Precision | Binades | Specials | Defined in |
|---|---|---|---|---|---|---|---|---|---|
| FP32 | 1·8·23 | 127 | 3.40282 × 10³⁸ | 2⁻¹²⁶ | 2⁻¹⁴⁹ | 24 bits (ulp at 1 = 2⁻²³) | 277 | ±∞ and NaN (all-ones exponent) | IEEE Standard for Floating-Point Arithmetic, Table 3.5, binary32 |
| FP16 | 1·5·10 | 15 | 65,504 | 2⁻¹⁴ | 2⁻²⁴ | 11 bits (ulp at 1 = 2⁻¹⁰) | 40 | ±∞ and NaN (all-ones exponent) | IEEE Standard for Floating-Point Arithmetic, Table 3.5, binary16 |
| BF16 | 1·8·7 | 127 | 3.38953 × 10³⁸ | 2⁻¹²⁶ | 2⁻¹³³ | 8 bits (ulp at 1 = 2⁻⁷) | 261 | ±∞ and NaN (all-ones exponent) | Google Cloud, 1 sign, 8 exponent, 7 mantissa bits; FP32's exponent |
| FP8 E5M2 | 1·5·2 | 15 | 57,344 | 2⁻¹⁴ | 2⁻¹⁶ | 3 bits (ulp at 1 = 2⁻²) | 32 | ±∞ and NaN (all-ones exponent) | OCP 8-bit Floating Point Specification (OFP8), Tables 1 and 2 |
| FP8 E4M3 | 1·4·3 | 7 | 448 | 2⁻⁶ | 2⁻⁹ | 4 bits (ulp at 1 = 2⁻³) | 18 | NaN only (S.1111.111); no ∞ | OCP 8-bit Floating Point Specification (OFP8), Tables 1 and 2 |
| FP6 E3M2 | 1·3·2 | 3 | 28 | 2⁻² | 2⁻⁴ | 3 bits (ulp at 1 = 2⁻²) | 9 | none: every code is a number | OCP Microscaling Formats (MX) Specification, Table 4 |
| FP6 E2M3 | 1·2·3 | 1 | 7.5 | 2⁰ | 2⁻³ | 4 bits (ulp at 1 = 2⁻³) | 6 | none: every code is a number | OCP Microscaling Formats (MX) Specification, Table 4 |
| FP4 E2M1 | 1·2·1 | 1 | 6 | 2⁰ | 2⁻¹ | 2 bits (ulp at 1 = 2⁻¹) | 4 | none: every code is a number | OCP Microscaling Formats (MX) Specification, Table 5 |
Binades: the powers of two from the smallest subnormal to the largest value, inclusive. For E4M3 and E5M2 this gives 18 and 32, as the OCP FP8 specification's Table 2 states.
MX block formats
A block of 32 elements shares one E8M0 scale X = 2k (k from −127 to 127; code 255 is NaN). The concrete formats of the OCP MX specification 1.0 (Table 1):
| Format | Element | Element bits | Block | Scale | Bits per value | Largest element power of two |
|---|---|---|---|---|---|---|
| MXFP8 (E4M3) | FP8 E4M3 | 8 | 32 | E8M0 | 8.25 | 28 |
| MXFP8 (E5M2) | FP8 E5M2 | 8 | 32 | E8M0 | 8.25 | 215 |
| MXFP6 (E3M2) | FP6 E3M2 | 6 | 32 | E8M0 | 6.25 | 24 |
| MXFP6 (E2M3) | FP6 E2M3 | 6 | 32 | E8M0 | 6.25 | 22 |
| MXFP4 (E2M1) | FP4 E2M1 | 4 | 32 | E8M0 | 4.25 | 22 |
| MXINT8 | INT8, 6 fraction bits (±1 63/64) | 8 | 32 | E8M0 | 8.25 | 20 |
NF4 (NormalFloat 4)
Not a floating-point format: a code book of 16 values derived from quantiles of the normal distribution, scaled per block of 64 by the block's absolute maximum (QLoRA). The values are bitsandbytes' table; a value is stored as the nearest code, ties to the lower one.
| Code | Value | Upper decision threshold |
|---|---|---|
| 0 | -1 | -0.8480964004993439 |
| 1 | -0.6961928009986877 | -0.6106329262256622 |
| 2 | -0.5250730514526367 | -0.4599952697753906 |
| 3 | -0.39491748809814453 | -0.33967943489551544 |
| 4 | -0.28444138169288635 | -0.23460740596055984 |
| 5 | -0.18477343022823334 | -0.13791173323988914 |
| 6 | -0.09105003625154495 | -0.045525018125772476 |
| 7 | 0 | 0.03979014977812767 |
| 8 | 0.07958029955625534 | 0.1202552504837513 |
| 9 | 0.16093020141124725 | 0.2035212516784668 |
| 10 | 0.24611230194568634 | 0.2920137718319893 |
| 11 | 0.33791524171829224 | 0.3893125355243683 |
| 12 | 0.44070982933044434 | 0.5016634166240692 |
| 13 | 0.5626170039176941 | 0.6427869200706482 |
| 14 | 0.7229568362236023 | 0.8614784181118011 |
| 15 | 1 | – |
Rounding modes
The five modes of IEEE 754-2019 (section 4.3) and stochastic rounding. Overflow follows section 7.4 (towards-zero modes give the largest finite value), or the OCP saturating mode when it is chosen; formats without infinity saturate or give NaN, as their specifications say.
| Mode | Rounds |
|---|---|
| roundTiesToEven | to the nearer neighbour; a tie goes to the one whose last mantissa bit is 0 (IEEE 754's default, the only mode OCP requires) |
| roundTiesToAway | to the nearer neighbour; a tie goes away from zero |
| roundTowardZero | always towards zero (truncation) |
| roundTowardPositive | always towards +∞ |
| roundTowardNegative | always towards −∞ |
| stochastic | up with probability equal to the distance above the lower neighbour, in ulps (not an IEEE mode; MX allows other modes) |
Sources
- Google Cloud, The bfloat16 numerical format
- bitsandbytes, get_4bit_type('nf4') and dQuantizeNF4 (commit 8336490)
- Micikevicius et al., FP8 Formats for Deep Learning (arXiv 2209.05433)
- IEEE Standard for Floating-Point Arithmetic, IEEE Std 754-2019
- ml_dtypes: NumPy dtype extensions for FP8, FP6, FP4 and bfloat16
- OCP Microscaling Formats (MX) Specification, Version 1.0 (September 2023)
- Rouhani et al., Microscaling Data Formats for Deep Learning (arXiv 2310.10537)
- OCP 8-bit Floating Point Specification (OFP8), Revision 1.0 (2023-12-01 corrected biases)
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs (arXiv 2305.14314)
8 formats, 6 MX formats, 6 rounding modes.