2. Range and precision
The two 16-bit formats split their bits differently, and that is the whole difference between them:
- The exponent bits set the range: how big and how small a number can be. bfloat16 keeps float32’s 8 exponent bits, so it has float32’s range. float16 has only 5, so it stops at 65504.
- The mantissa bits set the precision: how many numbers there are between one power of 2 and the next.
The same 128 mantissa patterns are used between every two powers of 2: from 1 to 2, from 2 to 4, from 256 to 512. So the gap between neighbors grows with the size of the number. Precision is relative: about 2 to 3 correct digits for bfloat16, wherever the number is.
A result that falls between two neighbors is rounded to the nearer one. That is all a computer can store.
Most numbers are not stored exactly, even simple ones. 0.1 is not a fraction with a power of 2 at the bottom (the denominator), so each format stores the nearest number it has: 0.100000001490 in float32, 0.099975585938 in float16 and 0.100097656250 in bfloat16.
Press 0.1 in the lab, then switch between the three formats and compare how far each one is from 0.1. Then press 256.75 in bfloat16 and in float16: float16 has more mantissa bits, so it keeps the 0.75.
The same rounding happens in training. A weight gets a small change at every step, and in a 16-bit format a small change to a weight near 1 can be lost completely in the rounding.
Go deeper Exactly halfway (ties), and the numbers below the smallest normal one
Ties. When a result lies exactly halfway between two neighbors, it goes to the neighbor whose last mantissa bit is 0 (“round half to even”). So 256 + 1 in bfloat16 is 256, not 258. Always rounding halves upward would make long sums slowly grow too big; this rule rounds up and down equally often.
Subnormals. When the exponent bits are all 0, the format drops the “1 +” and stores . These tiny numbers fill the space between 0 and the smallest normal number. In float16 the smallest one is . A smaller number rounds to whichever of 0 and is nearer, so anything below about becomes 0. When the exponent bits are all 1, the pattern means inf (mantissa 0) or nan (anything else).