Models & inference · Glossary term
What is Attention?
A mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using them to combine value vectors. Masks, position rules, or sparse patterns can restrict which positions participate.
“How a model focuses on important tokens.”
Why does Attention matter?
Attention lets a model route information between sequence positions, but it does not by itself explain or prove what the model understood.
What is the common confusion about Attention?
Attention weights are computation coefficients, not a faithful explanation of model reasoning.
Learn Attention in the course
Start with
- Self-Attention from Scratch
Attention is a lookup table where every word asks "who matters to me?" - and learns the answer. Implement scaled dot-product self-attention from scratch using only NumPy, including query/key/value…
Lessons that name Attention in a title or section
- Attention Mechanism — The Breakthrough
The decoder stops squinting at a compressed summary and starts looking at the whole source. Everything after this is attention plus engineering. Lesson 09 ended on a measured failure.
- Speech Recognition (ASR) — CTC, RNN-T, Attention
Speech recognition is audio classification at every timestep, glued together by a sequence model that knows English and silence. CTC, RNN-T, and attention are the three ways to do it.
- Multi-Head Attention
One attention head learns one relation at a time. Eight heads learn eight. Heads are free. Take more of them. A single self-attention head computes one attention matrix.
- KV Cache, Flash Attention & Inference Optimization
Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different tricks. A naive autoregressive decoder does O(N²) work to generate N tokens: at each step…
- Attention Variants — Sliding Window, Sparse, Differential
Full attention is a circle. Every token sees every token, and memory pays the price. Four variants bend the shape of the circle and recover half the cost.
- Differential Attention (V2)
Softmax attention spreads a small amount of probability over every non-matching token. Over 100k tokens that noise adds up and drowns the signal.
- Native Sparse Attention (DeepSeek NSA)
At 64k tokens, attention eats 70-80% of decode latency. Every open-model lab has a plan to fix it. DeepSeek's NSA (ACL 2025 best paper) is the one that stuck: three parallel attention branches —…
- Tensor Operations
Tensors are the common language between data and deep learning. Every image, every sentence, every gradient flows through them.
Taught in Phase 07: Transformers Deep Dive.
Also covered in Phase 01: Math Foundations, Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio, Phase 10: LLMs from Scratch, Phase 12: Multimodal AI and Phase 19: Capstone Projects.
Related terms
- Self-AttentionAttention in which queries, keys, and values are derived from the same sequence representation.
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
- Cross-AttentionAttention in which the query representation comes from one sequence or representation while keys and values come from another.
- FlashAttentionAn exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding…
- SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
- Visual GroundingConnecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.
Sources
More terms in Models & inference
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.