Skip to content
AI360Xpert
Beta
Core ML

Visual explainer

Tensor Processing Units

How TPUs trade general-purpose flexibility for massive, hardware-accelerated matrix multiplication using a systolic array.

A CPU fetches memory for every single math step.
A CPU fetches memory for every single math step.

You run matrix math on a CPU and most time goes to trips on every op. Each step fetches inputs, computes once, and writes back, so memory sets your speed limit.

One MAC Works

One MAC does multiply plus add as 2 times 3 plus 4 is 10.
One MAC does multiply plus add as 2 times 3 plus 4 is 10.

You build from one multiply-add cell that never moves data far. It takes 2 times 3 plus 4 and returns 10, and thousands of these cells form your grid.

Sums Flow Down

A 2 by 2 flow computes 19, 22, 43 and 50 together.
A 2 by 2 flow computes 19, 22, 43 and 50 together.

You pin weights [[5,6],[7,8]] and stream inputs [[1,2],[3,4]] across in one pass. Columns add as they fall and you read 19, 22, 43, 50 without middle fetches.

Weights Stay Put

Weights stay put while activations stream through each cell.
Weights stay put while activations stream through each cell.

You load each weight once and reuse it many times. Activations slide right, partial sums drop down, and only edges touch memory while the middle keeps computing.

Reuse Beats Refetch

Shared inputs cut fetches while outputs match exactly.
Shared inputs cut fetches while outputs match exactly.

You run the same 2 by 2 twice on one scale. Refetching grabs 8 inputs for 4 results, while flowing reuses rows and columns and still prints 19, 22, 43, 50.

Pick Dense Work

Pick TPU for dense math and CPU when branches dominate.
Pick TPU for dense math and CPU when branches dominate.

You choose this chip when dense multiplies fill your run. Flip back to CPU or GPU when branches, sparse zeros, or odd shapes leave most cells idle.

Where It Breaks

Sparse branching stalls the array and wastes most cells.
Sparse branching stalls the array and wastes most cells.

You feed sparse branching code with mostly zeros. Cells wait for decisions, neighbors sit idle, and your big grid runs slower than a small flexible chip.

The Quick Version

  • CPU fetches memory on every step.
  • One MAC returns 2 times 3 plus 4 as 10.
  • 2 by 2 flow prints 19, 22, 43, 50.
  • Weights stay, data streams through cells.
  • Reuse matches refetch with fewer trips.
  • Dense math here, branches elsewhere.
  • Sparse branches idle most of the grid.