Visual explainer
BERT vs GPT
One mask lets GPT predict the next word from the past, while no mask lets BERT fill a hidden word from both sides of the sentence.
Finishing a sentence and judging one are different jobs. The first needs strict left-to-right order because the future must stay unseen, while the second wants the whole sentence at once.
Past Predicts Next
You feed GPT cat, sat, on and it scores mat at 0.60, rug at 0.25 and moon at 0.15. Every guess comes from the past, never the future, which is exactly what training enforces.
Future Stays Hidden
The mask blanks the upper triangle, so row three spreads 0.665, 0.245 and 0.090 over past words only. Grey cells read nothing at all.
Both Sides Vote
You hide bank and let the full sentence vote. River and side agree, so bank lands at 0.80 with no left-to-right guessing involved.
Same Words, Different Wins
GPT completes the phrase at 0.60 confidence while BERT labels its tone at 0.90. Same scale, same words, and each wins its own job because each mask fits its question.
Match Reader to Job
Generating text? Take the masked decoder. Labeling or scoring? Take the full-view encoder. When you need both, stack them into an encoder-decoder.
Where It Breaks
An encoder has no next token slot, so asking it to write gives you a blank page with full understanding. Use the wrong reader and the job is the wrong shape entirely, no matter how big the model gets.
The Quick Version
- Finishing needs order; judging wants everything.
- GPT scores mat 0.60 from the past alone.
- The causal mask blanks the whole future.
- BERT fills the blank from both sides.
- Each reader wins only its own job.
- Generate with decoders, label with encoders.
- Encoders cannot write; that slot is missing.
What to Read Next
- TransformersProcessing five words at once through softmax weights and stacked masks, instead of squeezing a sentence through one thin recurrent pipe.
- Attention MechanismHow attention scores each input word, blends by weight, and rebuilds focus each step.
- Sequence-to-SequenceFold the input into one fixed vector with an encoder, then unroll it token by token with a decoder that feeds itself.