/learn
Systolic arrays, cycle by cycle
One chapter per idea, each opening with an animation. Every frame is a cycle of this site's model of the array (checked exactly against its Python reference, and against a SystemVerilog array in Verilator). Toggle layers (Concept / Maths / Code) inside any chapter to choose how deep to go. How a GPU does the same job is on GPU Kernels Explained; for slides, see the Google TPU series.
The memory wall, and Kung's answer: fetch each number once and pass it from processor to processor, so a grid of n × n multipliers needs only 2n words a cycle.
The TPU's dataflow: weights sit still, skewed activations march right, partial sums flow down and drain out of the bottom. Every register, every cycle.
The same matrix multiply in three dataflows side by side: what stays put, what moves, and what that costs in memory reads, link hops and register writes.
Why the inputs arrive as a staircase, how long the array takes to fill and empty, and what that does to utilisation as the matrices grow or stop fitting.
A matrix bigger than the array, cut into weight tiles and run back to back: double-buffered weights hide each tile's load, and the dataflow decides how many words the buffer must deliver.
One processing element, bit by bit: a three-stage MAC that multiplies bfloat16 or FP16 exactly and accumulates in FP32, checked register for register against the author's RTL.
- 07The TPU
The matrix units in a chip (MXU, vector and scalar units, on-chip memory, HBM) and the chips in a pod: an all-reduce on a 2-D torus, step by step.
GPU tensor cores against systolic arrays, Eyeriss's dataflow taxonomy with row-stationary animated, and computing in or near memory.
An ONNX model lowered onto the array: im2col turns a convolution into a GEMM, the GEMMs become weight tiles, and the cycle-accurate count meets a simulator's estimate.