Models & inference · Glossary term
What is Self-Attention?
Attention in which queries, keys, and values are derived from the same sequence representation. Scaled similarity scores are normalized and used to combine values, subject to causal, padding, local, or other masks.
“Tokens deciding which other tokens matter.”
Why does Self-Attention matter?
It builds context-sensitive token representations, but the permitted attention pattern depends on the architecture.
What is the common confusion about Self-Attention?
Not every token can always attend to every other token. Causal and sparse models intentionally restrict connections.
Learn Self-Attention in the course
Start with
- Self-Attention from Scratch
Attention is a lookup table where every word asks "who matters to me?" - and learns the answer. Implement scaled dot-product self-attention from scratch using only NumPy, including query/key/value…
Lessons that name Self-Attention in a title or section
- Multi-Head Self-Attention
One linear projection, three views, H parallel heads, one mask. The attention block as the model actually uses it. Implement a batched Query/Key/Value projection as a single linear layer split into…
- Pre-Training a Mini GPT (124M Parameters)
GPT-2 Small has 124 million parameters. That's 12 transformer layers, 12 attention heads, and 768-dimensional embeddings. You can train it from scratch on a single GPU in a few hours.
- Vision Transformer Encoder
Patches alone do not see. A 12-layer pre-LN transformer with 12 attention heads turns the sequence of patch tokens into a sequence of contextual tokens, with the CLS token pooling whole-image…
Taught in Phase 07: Transformers Deep Dive.
Also covered in Phase 10: LLMs from Scratch and Phase 19: Capstone Projects.
Related terms
- AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- Context WindowThe maximum token capacity available to one model inference under a specific model and API contract.
- Cross-AttentionAttention in which the query representation comes from one sequence or representation while keys and values come from another.
- FlashAttentionAn exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding…
- Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.