Visual explainer
How A Decision Head Works
Skip the word by word loop: read a request once with an encoder and answer with small typed heads that output choices, scores and flags.
Ask a chatbot to route tickets and it narrates JSON one token at a time. Twenty four tokens to say billing, and a single missing bracket kills the parse. You pay per token for the privilege of debugging punctuation.
Read It Once
An encoder instead reads everything in a single pass. Invoice 4821, charged twice, refund the second: all weighed together into one dense vector like 0.8, 0.2, 0.5, 0.1. Nothing is generated, nothing can hallucinate mid-sentence.
Three Heads Answer
Small specialist heads sit atop that vector and fire together. Choice says billing at 0.91, score rates priority 0.87, flag says no human at 0.94. Three typed answers land in the same instant the vector does.
Numbers You Audit
Probabilities beat prose for machines. Billing 0.91 plus shipping 0.06 plus returns 0.03 equals exactly 1.00, so you can threshold, log, and audit every call. Try doing that with a paragraph.
Proof It Worked
Same routing job, two workloads. Spelling the label costs 24 sequential steps that cannot parallelise. Choosing it costs one forward pass that answers everything at once. Fewer moving parts, fewer failures.
Router Or Writer
Pick by the answer shape. If every valid answer pre-exists as a label, heads win on speed and safety. If the reply needs fresh sentences, summaries, or code, only a generator qualifies. Different jobs, different tools.
Where It Breaks
Heads choose, they never compose. Hand them a summarise this thread request and every head abstains: there are no words to give. A router that writes is a generator in disguise, so hand writing jobs to the writer.
The Quick Version
- Spelling labels costs 24 risky tokens.
- Encoders read everything in one pass.
- Three heads answer simultaneously.
- Choice scores sum to 1.00 exactly.
- One pass replaces 24 steps.
- Closed answers suit heads, new words need writers.
- Heads choose but never compose.
What to Read Next
- Attention MechanismHow attention scores each input word, blends by weight, and rebuilds focus each step.
- TransformersProcessing five words at once through softmax weights and stacked masks, instead of squeezing a sentence through one thin recurrent pipe.
- The Agent Router ArchitectureHow a small classifier in front of big models sends each request to its cheapest lane.