Models & inference · Glossary term

What is Transformer?

A neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization. Encoder, decoder, and encoder-decoder variants use different masks and information flows.

What people say

“The architecture behind many modern language models.”

Why does Transformer matter?

Training can process many sequence positions in parallel, while autoregressive generation still produces outputs step by step.

What is the common confusion about Transformer?

Self-attention does not imply unrestricted all-to-all attention in every transformer.

Learn Transformer in the course

Start with

  • The Full Transformer — Encoder + Decoder

    Attention is the star. Everything else — residuals, normalization, feed-forward, cross-attention — is the scaffolding that lets you stack it deep.

    Phase 07: Transformers Deep Dive

Lessons that name Transformer in a title or section

  • Vision Transformers (ViT)

    Cut the image into patches, treat each patch as a word, run a standard transformer. Don't look back. Implement patch embedding, learned positional embedding, class token, and transformer encoder…

    Phase 04: Computer Vision

  • Diffusion Transformers & Rectified Flow

    The U-Net is not the secret of diffusion. Replace it with a transformer, swap the noise schedule for a straight-line flow, and suddenly you have SD3, FLUX, and every 2026 text-to-image model.

    Phase 04: Computer Vision

  • Text Generation Before Transformers — N-gram Language Models

    If a word is surprising, the model is bad. Perplexity makes surprise a number. Smoothing keeps it finite. Before transformers, before RNNs, before word embeddings, a language model predicted the…

    Phase 05: NLP: Foundations to Advanced

  • Why Transformers — The Problems with RNNs

    RNNs process tokens one at a time. Transformers process all tokens at once. That single architectural bet changed every scaling curve in deep learning after 2017.

    Phase 07: Transformers Deep Dive

  • Vision Transformers (ViT)

    An image is a grid of patches. A sentence is a grid of tokens. The same transformer eats both. Before 2020, computer vision meant convolutions.

    Phase 07: Transformers Deep Dive

  • Audio Transformers — Whisper Architecture

    Audio is an image of frequency over time. Whisper is a ViT that eats mel spectrograms and speaks back. Before Whisper (OpenAI, Radford et al.

    Phase 07: Transformers Deep Dive

  • Build a Transformer from Scratch — The Capstone

    Thirteen lessons. One model. No shortcuts. You've read every paper. You've implemented attention, multi-head splits, positional encodings, encoder and decoder blocks, BERT and GPT losses, MoE, KV…

    Phase 07: Transformers Deep Dive

  • Vision Transformers and the Patch-Token Primitive

    Before anything multimodal, an image has to become a sequence of tokens a transformer can eat. The 2020 ViT paper answered this with 16x16 pixel patches, a linear projection, and a position embedding.

    Phase 12: Multimodal AI

Taught in Phase 07: Transformers Deep Dive.

Also covered in Phase 01: Math Foundations, Phase 02: ML Fundamentals, Phase 03: Deep Learning Core, Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio, Phase 10: LLMs from Scratch, Phase 12: Multimodal AI and Phase 19: Capstone Projects.

  • AttentionA mechanism that forms contextual representations by comparing query vectors with key vectors, normalizing the resulting scores, and using…
  • Self-AttentionAttention in which queries, keys, and values are derived from the same sequence representation.
  • EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
  • DecoderA component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and…
  • GPTGenerative Pre-trained Transformer, a family label for generative transformer models pretrained on sequence-prediction objectives and…
  • Inductive BiasStructural or statistical assumptions that favor some functions or representations over others.
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • MoE (Mixture of Experts)An architecture with multiple expert subnetworks and a learned router that selects a subset for each input unit, often each token.
  • Multimodal ModelA model that learns from, relates, or generates more than one modality through representation, alignment, fusion, translation, or…
  • Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…

Sources

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.