Skip to content
AI360Xpert
Beta
Core ML

Visual explainer

How A Decision Head Works

Skip the word by word loop: read a request once with an encoder and answer with small typed heads that output choices, scores and flags.

Spelling out a routing label takes 24 tokens, and one broken bracket ruins it.
Spelling out a routing label takes 24 tokens, and one broken bracket ruins it.

Ask a chatbot to route tickets and it narrates JSON one token at a time. Twenty four tokens to say billing, and a single missing bracket kills the parse. You pay per token for the privilege of debugging punctuation.

Read It Once

An encoder reads the whole ticket once and compresses it into one vector.
An encoder reads the whole ticket once and compresses it into one vector.

An encoder instead reads everything in a single pass. Invoice 4821, charged twice, refund the second: all weighed together into one dense vector like 0.8, 0.2, 0.5, 0.1. Nothing is generated, nothing can hallucinate mid-sentence.

Three Heads Answer

Three small heads answer at once: billing at 0.91, priority 0.87, no human needed.
Three small heads answer at once: billing at 0.91, priority 0.87, no human needed.

Small specialist heads sit atop that vector and fire together. Choice says billing at 0.91, score rates priority 0.87, flag says no human at 0.94. Three typed answers land in the same instant the vector does.

Numbers You Audit

Choice scores sum to exactly 1.00, with billing holding 0.91 of it.
Choice scores sum to exactly 1.00, with billing holding 0.91 of it.

Probabilities beat prose for machines. Billing 0.91 plus shipping 0.06 plus returns 0.03 equals exactly 1.00, so you can threshold, log, and audit every call. Try doing that with a paragraph.

Proof It Worked

One forward pass replaces 24 decoding steps on the same job.
One forward pass replaces 24 decoding steps on the same job.

Same routing job, two workloads. Spelling the label costs 24 sequential steps that cannot parallelise. Choosing it costs one forward pass that answers everything at once. Fewer moving parts, fewer failures.

Router Or Writer

Route with heads when the answer is a choice, generate when words must be new.
Route with heads when the answer is a choice, generate when words must be new.

Pick by the answer shape. If every valid answer pre-exists as a label, heads win on speed and safety. If the reply needs fresh sentences, summaries, or code, only a generator qualifies. Different jobs, different tools.

Where It Breaks

Asked to summarize, the heads have no words to give and must sit out.
Asked to summarize, the heads have no words to give and must sit out.

Heads choose, they never compose. Hand them a summarise this thread request and every head abstains: there are no words to give. A router that writes is a generator in disguise, so hand writing jobs to the writer.

The Quick Version

  • Spelling labels costs 24 risky tokens.
  • Encoders read everything in one pass.
  • Three heads answer simultaneously.
  • Choice scores sum to 1.00 exactly.
  • One pass replaces 24 steps.
  • Closed answers suit heads, new words need writers.
  • Heads choose but never compose.