Visual explainer
Transformers
Processing five words at once through softmax weights and stacked masks, instead of squeezing a sentence through one thin recurrent pipe.
An RNN reads left to right and files every word into one fixed vector. By word thirty the early words have faded, and your decoder guesses from mush it cannot check.
What Scores What
You compare one query against every key with a dot product. Bank meets river at 8 and cash at 1, so river wins the blend before any learning happens.
Shares Must Sum to One
Raw dots 16, 8 and 0 get divided by 8, then softmax turns 2, 1 and 0 into 0.665, 0.245 and 0.090. The output is that exact blend of the values.
Two Heads Beat One
One head tracks wording and parks 0.70 on river; the other tracks order and parks 0.60 on the last word. Concatenated, bank finally means river bank.
Five Steps Become One
The same five words cost an RNN five serial steps and 50 milliseconds. Attention scores all pairs at once for 12 milliseconds, and length stops mattering.
Mask Fits the Job
Writing the next word? Hide the future with a causal mask. Labeling tone? Let every word see all. When you need both, stack an encoder under a decoder.
Where It Breaks
Pairs grow with length squared, so 4,096 tokens need 16.8 million pairs and 64 times the memory of 512. Past that, plain attention stops being affordable.
The Quick Version
- One fixed vector overflows on long sentences.
- Each query scores every key with a dot product.
- Softmax turns 2, 1, 0 into 0.665, 0.245, 0.090.
- Heads split wording and order, then concatenate.
- Five parallel steps replace five serial ones.
- Generate with a causal mask, label with a full one.
- Pairs grow with length squared, so budget accordingly.
What to Read Next
- Attention MechanismHow attention scores each input word, blends by weight, and rebuilds focus each step.
- BERT vs GPTOne mask lets GPT predict the next word from the past, while no mask lets BERT fill a hidden word from both sides of the sentence.
- Sequence-to-SequenceFold the input into one fixed vector with an encoder, then unroll it token by token with a decoder that feeds itself.