Inside a PE
One processing element, bit by bit: a three-stage MAC that multiplies bfloat16 or FP16 exactly and accumulates in FP32, checked register for register against the author's RTL.
Loading the animation…
Concept
Every chapter so far drew a PE as a box that multiplies and adds in one cycle. Inside, it is a small pipeline, and this animation is one: the author's mixed-precision MAC unit, a SystemVerilog design with three stages, run cycle by cycle by this site's model of it.
- Stage 1 latches the two operands, (an activation) and (a weight), with a
validbit and aclearbit that marks the first term of a new dot product. - Stage 2 multiplies them. For two bfloat16 numbers the 8-bit significands (7 stored bits plus the hidden 1) multiply into at most 16 bits, which FP32's 24-bit significand holds exactly, so the product needs no rounding. FP16's 11-bit significands give 22 bits: still exact. INT8 × INT8 fits 16 bits.
- Stage 3 adds the product to the accumulator. This is the only place the PE rounds: the smaller of the two numbers is shifted right to line up the binary points, the sum is rounded back to 24 significant bits (to nearest, ties to even), and anything that would be subnormal is flushed to zero. In INT8 mode the accumulator is a 32-bit integer that saturates instead of wrapping.
Each colour band in the picture is a field of the number: sign (purple), exponent (sky), mantissa (green); a lit cell is a 1 bit. The two dot products of 4 terms run back to back, and the second starts on the cycle the clear bit reaches stage 3: no bubble between them. In FP16 mode, 2 of the demonstration's adds round; the caption shows the exact sum next to the stored one whenever they differ.
The pipeline is why the arrays in chapters 2 to 5 can run at a high clock rate: each stage does a third of the work. It adds latency (a product reaches the accumulator three cycles after its operands arrive) but not cost per MAC, because a new pair enters every cycle.
Concept
Why bfloat16 in and FP32 out. bfloat16 is the top half of an FP32: the same 8-bit exponent, so the same range, but only 7 stored mantissa bits. Google's TPUs multiply in bfloat16 and accumulate in FP32 by default (Google Cloud TPU documentation), and Kalamkar et al. (2019) found that training with bfloat16 tensors reached the same results as FP32 across image, speech, language and recommendation models, in the same number of iterations and with no change of hyperparameters. The narrow mantissa makes the multiplier small (Google's bfloat16 article notes that a multiplier's size scales with the square of the mantissa width) while the wide accumulator keeps the long sums of a matrix multiply accurate. Why a long sum needs the extra bits, and what it does to the error, is chapter 3 of Numerics Explained; how formats from INT8 to FP4 map onto hardware is its chapter 10.
Checked against the RTL. The site's Python model of this PE (reference/pe.py) is checked against the RTL itself: rtl/check_rtl.py builds the vendored design in Verilator 5.020, drives 5 sequences (1,216 operations, FP16 and INT8, with idle cycles and exponents across FP16's whole normal range), and compares every pipeline register after every clock edge: 11,079 values, 0 mismatches. The bfloat16 mode is this site's addition on the same pipeline (the RTL has FP16 and INT8 modes); it is checked against NumPy's float32 arithmetic instead.
Maths
A bfloat16 number with sign , biased exponent and mantissa (7 bits) is . The product of two is : the significand product is an integer below , so it is exactly representable in FP32 (24 bits) as long as the exponent stays in range. The adder then computes
the exact sum rounded once. The model computes the sum in double precision (53 bits) and rounds that to FP32; for addition this double rounding cannot change the result, because (Figueroa, 1995). The RTL gets the same answer from three extra bits below the significand (guard, round and a sticky OR of everything shifted out), and the comparison above confirms it on every add.
Code
The rounding step of the RTL's FP32 adder, from rtl/mac_unit.sv (vendored unchanged from the author's Interview_RTL_LLM_Accelerators repository):
// Round to nearest, ties to even: guard = norm_sig[2],
// round|sticky = norm_sig[1:0], LSB of the kept mantissa = norm_sig[3]
round_up = norm_sig[2] & ((|norm_sig[1:0]) | norm_sig[3]);
rounded_sig = norm_sig[26:3] + {23'd0, round_up};
Stage 3's choice between starting a new dot product and adding to the old one:
if (s2_clear && s2_mode) begin
// Clear then accumulate (or just clear to +0.0 if no valid data)
fp32_acc_next = s2_valid ? s2_fp32_product : 32'h00000000;
end else if (s2_valid && s2_mode) begin
fp32_acc_next = fp32_add_result;
The same step in the site's TypeScript model, cut from src/lib/sa/pe.ts:
export function fp32Add(acc: number, p: number): number {
const r = f32Bits(f32Value(acc) + f32Value(p));
const v = f32Value(r);
if (v === 0 || Math.abs(v) < FP32_MIN_NORMAL) return 0;
return r;
}