/model
The cycle-accurate model
The model computes C = AB (A is M × K, B is K × N) on an R × C array. A frame is the state of every register after the clock edge that ends a cycle, computed only from the previous frame and the values entering at the edges, so data moves exactly one PE per cycle. A PE holds up to four registers: s, the stationary operand; h, the operand moving right; v, what moves down (a partial sum, or in output-stationary an operand); and acc, output-stationary's accumulator. Each register holds a tag naming the element of A, B or C it carries, so the captions can say which value is where. Source: reference/systolic.py and its exact TypeScript port src/lib/sa/model.ts.
Dataflows and closed forms
The array block each dataflow needs, what stays in it, where and when each product happens, the cycles of one pass (load, compute and drain) and the partial sums moved between PEs. The tests count all of these from the frames.
| Dataflow | Block | Stationary | Product | Cycles | Partial-sum hops |
|---|---|---|---|---|---|
| weight-stationary | K × N | B (weights) | aₘₖ in PE(k, n) at m + k + n | K + (M + K + N − 2) | M N (K − 1) |
| output-stationary | M × N | C (outputs) | aₘₖ, bₖₙ in PE(m, n) at k + m + n | (K + M + N − 2) + M | N M (M − 1) / 2 |
| input-stationary | K × M | A (inputs) | bₖₙ in PE(k, m) at n + k + m | K + (N + K + M − 2) | M N (K − 1) |
For a 128 × 128 × 128 problem on a 128 × 128 array:
| Dataflow | Cycles | Buffer reads | Buffer writes | PE-to-PE hops | Partial-sum hops | Register writes |
|---|---|---|---|---|---|---|
| weight-stationary | 510 | 32,768 | 16,384 | 5,201,920 | 2,080,768 | 5,251,072 |
| output-stationary | 510 | 32,768 | 16,384 | 5,201,920 | 1,040,384 | 7,331,840 |
| input-stationary | 510 | 32,768 | 16,384 | 5,201,920 | 2,080,768 | 5,251,072 |
The RTL cross-check
The design is a parameterised N × N output-stationary array with internal skew registers, from the author's Interview RTL challenge (modules pe and systolic_array, vendored unchanged at commit b37984f). rtl/tb_trace.sv drives one matrix multiply as the challenge's own testbench does and prints every accumulator after every rising edge; rtl/check_rtl.py builds it in Verilator 5.020, runs it and compares. Edge p is the model's compute cycle p − 1; after the last product (edge K + 2(N − 1)) the RTL holds its results and pulses valid_out, so there the check requires C = AB. Run with --shift 1, comparing each edge with the next cycle, it reports mismatches in every case: the comparison can fail. The same testbench in Vivado xsim 2025.2 gives byte-identical traces (run locally; CI has no Vivado).
| Case | N | K | Values in | Edges | Accumulators compared | Mismatches | valid_out at edge |
|---|---|---|---|---|---|---|---|
| n4_k4_small | 4 | 4 | -4 … 5 | 14 | 224 | 0 | 10 |
| n4_k7_int8 | 4 | 7 | -128 … 127 | 17 | 272 | 0 | 13 |
| n3_k5_small | 3 | 5 | -4 … 5 | 13 | 117 | 0 | 9 |
The weight- and input-stationary dataflows are not checked against RTL here; they are checked against a direct matrix multiply, their textbook timing and their closed forms.
The processing element against its RTL
Chapter 6's PE is the author's three-stage mixed-precision MAC unit (modules fp16_mul_to_fp32, fp32_add and mac_unit_mixed_precision of the MAC unit challenge, vendored unchanged at commit b37984f as rtl/mac_unit.sv). rtl/tb_mac_trace.sv drives one operation per cycle and prints every pipeline register after every rising edge; the same script compares them with reference/pe.py, which models the unit register for register (FP16 mode: the stage-1 operands, the exact FP32 product, the FP32 accumulator, valid and clear; INT8 mode: the operand bytes, the INT16 product and the INT32 accumulator). With --shift 1 every case reports mismatches. The same five runs in Vivado xsim 2025.2 give byte-identical traces (run locally). The site's bfloat16 mode has no RTL counterpart and is checked against NumPy.
| Case | Mode | Operations | Valid | Dot products | Edges | Registers compared | Mismatches |
|---|---|---|---|---|---|---|---|
| mac_fp16_demo | FP16 | 8 | 8 | 2 | 11 | 99 | 0 |
| mac_int8_demo | INT8 | 8 | 8 | 2 | 11 | 99 | 0 |
| mac_fp16_wide | FP16 | 400 | 331 | 25 | 403 | 3,627 | 0 |
| mac_fp16_long | FP16 | 400 | 316 | 4 | 403 | 3,627 | 0 |
| mac_int8_rand | INT8 | 400 | 319 | 25 | 403 | 3,627 | 0 |
Tiles back to back
Chapter 5 adds a shadow weight register to every PE and runs a GEMM's weight-stationary tiles back to back (simulateWsStream). The model checks, every cycle, that a PE swaps in the weight of the tile whose activation has just arrived and that every partial sum meets its next product in step; moving one tile's load or stream a cycle earlier makes it stop with an error. Its cycle counts equal the closed form Sj+1 = max(Sj + M, Lj+1 + kj+1), and with the shadow registers off they equal chapter 4's tiles in sequence.
The lowering against Torch_Sim_Frontend
Chapter 9 lowers a small ONNX model onto the array. scripts/check_simfront.py runs the author's Torch_Sim_Frontend (commit de3acd4) on the same file: its ONNX front end and GEMM rule give the same shapes, and its cycle-approximate array formula the same numbers, on a 8 × 8 array.
| Node | simfront category | GEMM (batch, M, K, N) | simfront cycles | Agree |
|---|---|---|---|---|
| Conv | matmul | 1 × 36 × 27 × 8 | 151 | yes |
| Relu | elementwise | none | – | yes |
| Flatten | view | none | – | yes |
| Gemm | matmul | 1 × 1 × 288 × 10 | 592 | yes |
Row-stationary, as a reference
For a convolution, Eyeriss's row-stationary dataflow gives PE(i, j) filter row i and input row i + j; each PE slides its filter row along its input row, one MAC per cycle, and each column of PEs adds its rows into one output row. The model runs it too (simulateRs): a 6 × 7 input and a 3 × 3 filter take 17 cycles on 3 × 4 PEs, and the output equals a direct convolution. Chapter 8 animates it.