Positional Encoding
A mechanism used in transformer networks to inject information about the absolute or relative order of tokens into their dense vector representations.
Think of It Like This
Like numbering the pages of a scattered manuscript so the reader inherently understands the chronological flow of the story.
Because standard attention operations process all tokens simultaneously without any inherent concept of sequence, positional encodings are mandatory. Original transformers used absolute sinusoidal waves added directly to the embeddings. Modern architectures largely utilize relative mechanisms like Rotary Position Embeddings (RoPE) or ALiBi.