numerics-explained
← /learn · 02

Rounding

Round to nearest, ties to even, against stochastic rounding: one is the most accurate on each value, the other the only one that is right on average.

Loading the animation…

Concept

Almost no result of arithmetic is exactly representable: it lands between two neighbouring values of the format, x↓x_\downarrow below and x↑x_\uparrow above, and something has to choose. IEEE 754's default, round to nearest, ties to even, takes the nearer neighbour, and when the result is exactly halfway takes the one whose last mantissa bit is 0. Its error is never more than half an ulp, the best any rounding can promise for a single value. The rule for ties matters: always rounding halves up would push every tie the same way. IEEE 754-2019 defines four other modes (ties away from zero, and towards +∞+\infty, −∞-\infty and zero); the formats page lists them, and the library implements all five.

Stochastic rounding flips a biased coin instead. If xx is 0.3 of the way from x↓x_\downarrow to x↑x_\uparrow, it rounds up with probability 0.3 and down with probability 0.7. Each single result is worse (the error can approach a whole ulp), but on average it is exact: the expected value of the rounded number is xx itself.

The animation feeds the same inputs to both. With inputs that sit just above a representable value, nearest rounding always goes down, and its errors pile up on one side: after 64 inputs its mean error is −0.1772 ulp, while stochastic rounding's is −0.005308 ulp. Switch to inputs spread anywhere and both means stay near zero; switch to exact ties and watch nearest-even alternate by the last bit.

Loading the animation…

Concept

Bias matters most when many small changes are added to something large, which is what training does: each step adds a tiny update to every weight. The second animation adds δ\delta, an eighth of an ulp, to h=1h = 1 128 times in FP16. The exact answer is 1.015625. Rounding to nearest gives 1: every single addition lands closer to hh than to the next value, so it is rounded straight back and nothing ever accumulates. Stochastic rounding moves up a whole ulp about one time in eight and ends at 1.0185546875.

This is why low-precision training either keeps an FP32 "master" copy of the weights to add the updates to (the mixed-precision recipe of Micikevicius et al.) or rounds stochastically (Gupta et al. trained networks with 16-bit fixed-point numbers this way). Inference has no updates to lose, and uses round to nearest even.

Maths

Let xx lie in [x↓,x↑][x_\downarrow, x_\uparrow] with x↑−x↓=ulpx_\uparrow - x_\downarrow = \mathrm{ulp} and fraction f=(x−x↓)/ulp∈[0,1)f = (x - x_\downarrow)/\mathrm{ulp} \in [0, 1). Stochastic rounding returns x↑x_\uparrow with probability ff, so

E[SR(x)]=x↓+f ulp=x,Var⁡[SR(x)]=f(1−f) ulp2≤14ulp2.\mathbb{E}[\mathrm{SR}(x)] = x_\downarrow + f\,\mathrm{ulp} = x, \qquad \operatorname{Var}[\mathrm{SR}(x)] = f(1-f)\,\mathrm{ulp}^2 \le \tfrac{1}{4}\mathrm{ulp}^2 .

Round to nearest has error min⁡(f,1−f) ulp≤12ulp\min(f, 1-f)\,\mathrm{ulp} \le \tfrac{1}{2}\mathrm{ulp} but no randomness to average away: for inputs with f<12f < \tfrac{1}{2} every error is negative. In the stagnation example h+δh + \delta has f=18f = \tfrac{1}{8}, so nearest returns hh every time, while the stochastic sum after tt steps has mean 1+tδ1 + t\delta and standard deviation at most 12t ulp\tfrac{1}{2}\sqrt{t}\,\mathrm{ulp} while it stays in one binade.

The library draws the coin from a 32-bit integer uu: it rounds up when u/232<fu / 2^{32} < f.

Code

The rounding decision, cut from src/lib/num/model.ts (frac is ff; n is the magnitude in ulps, rounded down):

switch (mode) {
  case "rne":
    return frac > 0.5 || (frac === 0.5 && n % 2 === 1);
  case "rna":
    return frac >= 0.5;
  case "rtz":
    return false;
  case "rup":
    return !neg;
  case "rdn":
    return neg;
  case "sr":
    return u / TWO32 < frac;
}