Models & inference · Glossary term

What is Attention?

A mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using them to combine value vectors. Masks, position rules, or sparse patterns can restrict which positions participate.

What people say

“How a model focuses on important tokens.”

Why does Attention matter?

Attention lets a model route information between sequence positions, but it does not by itself explain or prove what the model understood.

What is the common confusion about Attention?

Attention weights are computation coefficients, not a faithful explanation of model reasoning.

Learn Attention in the course

Start with

  • Self-Attention from Scratch

    Attention is a lookup table where every word asks "who matters to me?" - and learns the answer. Implement scaled dot-product self-attention from scratch using only NumPy, including query/key/value…

    Phase 07: Transformers Deep Dive

Lessons that name Attention in a title or section

  • Attention Mechanism — The Breakthrough

    The decoder stops squinting at a compressed summary and starts looking at the whole source. Everything after this is attention plus engineering. Lesson 09 ended on a measured failure.

    Phase 05: NLP: Foundations to Advanced

  • Speech Recognition (ASR) — CTC, RNN-T, Attention

    Speech recognition is audio classification at every timestep, glued together by a sequence model that knows English and silence. CTC, RNN-T, and attention are the three ways to do it.

    Phase 06: Speech & Audio

  • Multi-Head Attention

    One attention head learns one relation at a time. Eight heads learn eight. Heads are free. Take more of them. A single self-attention head computes one attention matrix.

    Phase 07: Transformers Deep Dive

  • KV Cache, Flash Attention & Inference Optimization

    Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different tricks. A naive autoregressive decoder does O(N²) work to generate N tokens: at each step…

    Phase 07: Transformers Deep Dive

  • Attention Variants — Sliding Window, Sparse, Differential

    Full attention is a circle. Every token sees every token, and memory pays the price. Four variants bend the shape of the circle and recover half the cost.

    Phase 07: Transformers Deep Dive

  • Differential Attention (V2)

    Softmax attention spreads a small amount of probability over every non-matching token. Over 100k tokens that noise adds up and drowns the signal.

    Phase 10: LLMs from Scratch

  • Native Sparse Attention (DeepSeek NSA)

    At 64k tokens, attention eats 70-80% of decode latency. Every open-model lab has a plan to fix it. DeepSeek's NSA (ACL 2025 best paper) is the one that stuck: three parallel attention branches —…

    Phase 10: LLMs from Scratch

  • Tensor Operations

    Tensors are the common language between data and deep learning. Every image, every sentence, every gradient flows through them.

    Phase 01: Math Foundations

Taught in Phase 07: Transformers Deep Dive.

Also covered in Phase 01: Math Foundations, Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio, Phase 10: LLMs from Scratch, Phase 12: Multimodal AI and Phase 19: Capstone Projects.

  • Self-AttentionAttention in which queries, keys, and values are derived from the same sequence representation.
  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • Cross-AttentionAttention in which the query representation comes from one sequence or representation while keys and values come from another.
  • FlashAttentionAn exact attention algorithm that tiles the computation to reduce transfers between accelerator memory levels while avoiding…
  • SoftmaxA function defined by `softmax(x_i) = exp(x_i) / sum(exp(x_j))`, implemented with numerical stabilization.
  • Visual GroundingConnecting a language expression to spatial evidence in an image or video, such as a region, object, mask, or tracked entity.

Sources

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.