Models & inference · Glossary term

What is Self-Attention?

Attention in which queries, keys, and values are derived from the same sequence representation. Scaled similarity scores are normalized and used to combine values, subject to causal, padding, local, or other masks.

What people say

“Tokens deciding which other tokens matter.”

Why does Self-Attention matter?

It builds context-sensitive token representations, but the permitted attention pattern depends on the architecture.

What is the common confusion about Self-Attention?

Not every token can always attend to every other token. Causal and sparse models intentionally restrict connections.

Learn Self-Attention in the course

Start with

  • Self-Attention from Scratch

    Attention is a lookup table where every word asks "who matters to me?" - and learns the answer. Implement scaled dot-product self-attention from scratch using only NumPy, including query/key/value…

    Phase 07: Transformers Deep Dive

Lessons that name Self-Attention in a title or section

  • Multi-Head Self-Attention

    One linear projection, three views, H parallel heads, one mask. The attention block as the model actually uses it. Implement a batched Query/Key/Value projection as a single linear layer split into…

    Phase 19: Capstone Projects

  • Pre-Training a Mini GPT (124M Parameters)

    GPT-2 Small has 124 million parameters. That's 12 transformer layers, 12 attention heads, and 768-dimensional embeddings. You can train it from scratch on a single GPU in a few hours.

    Phase 10: LLMs from Scratch

  • Vision Transformer Encoder

    Patches alone do not see. A 12-layer pre-LN transformer with 12 attention heads turns the sequence of patch tokens into a sequence of contextual tokens, with the CLS token pooling whole-image…

    Phase 19: Capstone Projects

Taught in Phase 07: Transformers Deep Dive.

Also covered in Phase 10: LLMs from Scratch and Phase 19: Capstone Projects.

  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • Context WindowThe maximum token capacity available to one model inference under a specific model and API contract.
  • Cross-AttentionAttention in which the query representation comes from one sequence or representation while keys and values come from another.
  • FlashAttentionAn exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding…
  • Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.