Models & inference · Glossary term

What is Encoder?

A component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any masks, so each position can incorporate context from across the input.

What people say

“The input side of a model.”

What is the common confusion about Encoder?

Encoder-only models can produce outputs through task heads even though they are not typically used for autoregressive text generation.

Learn Encoder in the course

Lessons that name Encoder in a title or section

  • The Full Transformer — Encoder + Decoder

    Attention is the star. Everything else — residuals, normalization, feed-forward, cross-attention — is the scaffolding that lets you stack it deep.

    Phase 07: Transformers Deep Dive

  • Janus-Pro: Decoupled Encoders for Unified Multimodal Models

    Unified multimodal models have an unavoidable tension. Understanding wants semantic features — SigLIP or DINOv2 output vectors rich with concept-level information.

    Phase 12: Multimodal AI

  • Vision Encoder Patches

    A vision model that reads pixels needs a tokenizer for pixels. Patch embedding is that tokenizer. Cut the image into a grid of squares, flatten each square, project it through one linear layer, then…

    Phase 19: Capstone Projects

  • Vision Transformer Encoder

    Patches alone do not see. A 12-layer pre-LN transformer with 12 attention heads turns the sequence of patch tokens into a sequence of contextual tokens, with the CLS token pooling whole-image…

    Phase 19: Capstone Projects

  • Semantic Segmentation — U-Net

    Segmentation is classification at every pixel. U-Net makes it work by pairing a downsampling encoder with an upsampling decoder and wiring skip connections between them.

    Phase 04: Computer Vision

  • Vision Transformers (ViT)

    Cut the image into patches, treat each patch as a word, run a standard transformer. Don't look back. Implement patch embedding, learned positional embedding, class token, and transformer encoder…

    Phase 04: Computer Vision

  • Diffusion Transformers & Rectified Flow

    The U-Net is not the secret of diffusion. Replace it with a transformer, swap the noise schedule for a straight-line flow, and suddenly you have SD3, FLUX, and every 2026 text-to-image model.

    Phase 04: Computer Vision

  • Sequence-to-Sequence Models

    Two RNNs pretending to be a translator. The bottleneck they hit is the reason attention exists. Classification maps a variable-length sequence to a single label.

    Phase 05: NLP: Foundations to Advanced

Covered in Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 07: Transformers Deep Dive, Phase 08: Generative AI, Phase 12: Multimodal AI and Phase 19: Capstone Projects.

  • DecoderA component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and…
  • TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
  • EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
  • Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…
  • VAE (Variational Autoencoder)A latent-variable model trained with a reconstruction objective and a regularization term that keeps an approximate posterior close to a…
  • Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.