Systolic Arrays Explained
How the hardware multiplies matrices
A matrix unit is a grid of multiply-accumulate cells that pass numbers to their neighbours every clock cycle: each operand is fetched from memory once and used by a whole row or column of cells. This site takes that machine apart one cycle at a time: which operand stays put, how the inputs are skewed so everything meets at the right moment, and where the cycles go.
Every animation is drawn from a cycle-accurate model of the array (Python reference, exact TypeScript port). Its output-stationary frames match a SystemVerilog array simulated in Verilator on every accumulator after every clock edge: 613 values, 0 mismatches.
Loading the animation…
01
Why systolic?
The memory wall, and Kung's answer: fetch each number once and pass it from processor to processor, so a grid of n × n multipliers needs only 2n words a cycle.
02
Weight-stationary, cycle by cycle
The TPU's dataflow: weights sit still, skewed activations march right, partial sums flow down and drain out of the bottom. Every register, every cycle.
03
Output- and input-stationary
The same matrix multiply in three dataflows side by side: what stays put, what moves, and what that costs in memory reads, link hops and register writes.
04
Skew, fill and drain
Why the inputs arrive as a staircase, how long the array takes to fill and empty, and what that does to utilisation as the matrices grow or stop fitting.
05
Tiling big GEMMs
A matrix bigger than the array, cut into weight tiles and run back to back: double-buffered weights hide each tile's load, and the dataflow decides how many words the buffer must deliver.
06
Inside a PE
One processing element, bit by bit: a three-stage MAC that multiplies bfloat16 or FP16 exactly and accumulates in FP32, checked register for register against the author's RTL.
07
The TPU
The matrix units in a chip (MXU, vector and scalar units, on-chip memory, HBM) and the chips in a pod: an all-reduce on a 2-D torus, step by step.
08
Other ways to build it
GPU tensor cores against systolic arrays, Eyeriss's dataflow taxonomy with row-stationary animated, and computing in or near memory.
09
From graph to silicon
An ONNX model lowered onto the array: im2col turns a convolution into a GEMM, the GEMMs become weight tiles, and the cycle-accurate count meets a simulator's estimate.
The model, its conventions and the RTL cross-check: the model.
Part of a family of companion sites: the Transformer Decoder Explainer (one forward pass), LLM Inference Explained (serving it), LLM Architectures Explained (how the models differ), GPU Kernels Explained (how a GPU runs the maths) and Numerics Explained (the number formats). This site is the silicon underneath. How it was built, and how to check it: about.