Rounding
Round to nearest, ties to even, against stochastic rounding: one is the most accurate on each value, the other the only one that is right on average.
Loading the animation…
Concept
Almost no result of arithmetic is exactly representable: it lands between two neighbouring values of the format, below and above, and something has to choose. IEEE 754's default, round to nearest, ties to even, takes the nearer neighbour, and when the result is exactly halfway takes the one whose last mantissa bit is 0. Its error is never more than half an ulp, the best any rounding can promise for a single value. The rule for ties matters: always rounding halves up would push every tie the same way. IEEE 754-2019 defines four other modes (ties away from zero, and towards , and zero); the formats page lists them, and the library implements all five.
Stochastic rounding flips a biased coin instead. If is 0.3 of the way from to , it rounds up with probability 0.3 and down with probability 0.7. Each single result is worse (the error can approach a whole ulp), but on average it is exact: the expected value of the rounded number is itself.
The animation feeds the same inputs to both. With inputs that sit just above a representable value, nearest rounding always goes down, and its errors pile up on one side: after 64 inputs its mean error is −0.1772 ulp, while stochastic rounding's is −0.005308 ulp. Switch to inputs spread anywhere and both means stay near zero; switch to exact ties and watch nearest-even alternate by the last bit.
Loading the animation…
Concept
Bias matters most when many small changes are added to something large, which is what training does: each step adds a tiny update to every weight. The second animation adds , an eighth of an ulp, to 128 times in FP16. The exact answer is 1.015625. Rounding to nearest gives 1: every single addition lands closer to than to the next value, so it is rounded straight back and nothing ever accumulates. Stochastic rounding moves up a whole ulp about one time in eight and ends at 1.0185546875.
This is why low-precision training either keeps an FP32 "master" copy of the weights to add the updates to (the mixed-precision recipe of Micikevicius et al.) or rounds stochastically (Gupta et al. trained networks with 16-bit fixed-point numbers this way). Inference has no updates to lose, and uses round to nearest even.
Maths
Let lie in with and fraction . Stochastic rounding returns with probability , so
Round to nearest has error but no randomness to average away: for inputs with every error is negative. In the stagnation example has , so nearest returns every time, while the stochastic sum after steps has mean and standard deviation at most while it stays in one binade.
The library draws the coin from a 32-bit integer : it rounds up when .
Code
The rounding decision, cut from src/lib/num/model.ts (frac is ; n is the magnitude in ulps, rounded down):
switch (mode) {
case "rne":
return frac > 0.5 || (frac === 0.5 && n % 2 === 1);
case "rna":
return frac >= 0.5;
case "rtz":
return false;
case "rup":
return !neg;
case "rdn":
return neg;
case "sr":
return u / TWO32 < frac;
}