Visual explainer
Tensor Processing Units
How TPUs trade general-purpose flexibility for massive, hardware-accelerated matrix multiplication using a systolic array.
You run matrix math on a CPU and most time goes to trips on every op. Each step fetches inputs, computes once, and writes back, so memory sets your speed limit.
One MAC Works
You build from one multiply-add cell that never moves data far. It takes 2 times 3 plus 4 and returns 10, and thousands of these cells form your grid.
Sums Flow Down
You pin weights [[5,6],[7,8]] and stream inputs [[1,2],[3,4]] across in one pass. Columns add as they fall and you read 19, 22, 43, 50 without middle fetches.
Weights Stay Put
You load each weight once and reuse it many times. Activations slide right, partial sums drop down, and only edges touch memory while the middle keeps computing.
Reuse Beats Refetch
You run the same 2 by 2 twice on one scale. Refetching grabs 8 inputs for 4 results, while flowing reuses rows and columns and still prints 19, 22, 43, 50.
Pick Dense Work
You choose this chip when dense multiplies fill your run. Flip back to CPU or GPU when branches, sparse zeros, or odd shapes leave most cells idle.
Where It Breaks
You feed sparse branching code with mostly zeros. Cells wait for decisions, neighbors sit idle, and your big grid runs slower than a small flexible chip.
The Quick Version
- CPU fetches memory on every step.
- One MAC returns 2 times 3 plus 4 as 10.
- 2 by 2 flow prints 19, 22, 43, 50.
- Weights stay, data streams through cells.
- Reuse matches refetch with fewer trips.
- Dense math here, branches elsewhere.
- Sparse branches idle most of the grid.