Models & inference · Glossary term
What is Encoder?
A component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any masks, so each position can incorporate context from across the input.
“The input side of a model.”
What is the common confusion about Encoder?
Encoder-only models can produce outputs through task heads even though they are not typically used for autoregressive text generation.
Learn Encoder in the course
Lessons that name Encoder in a title or section
- The Full Transformer — Encoder + Decoder
Attention is the star. Everything else — residuals, normalization, feed-forward, cross-attention — is the scaffolding that lets you stack it deep.
- Janus-Pro: Decoupled Encoders for Unified Multimodal Models
Unified multimodal models have an unavoidable tension. Understanding wants semantic features — SigLIP or DINOv2 output vectors rich with concept-level information.
- Vision Encoder Patches
A vision model that reads pixels needs a tokenizer for pixels. Patch embedding is that tokenizer. Cut the image into a grid of squares, flatten each square, project it through one linear layer, then…
- Vision Transformer Encoder
Patches alone do not see. A 12-layer pre-LN transformer with 12 attention heads turns the sequence of patch tokens into a sequence of contextual tokens, with the CLS token pooling whole-image…
- Semantic Segmentation — U-Net
Segmentation is classification at every pixel. U-Net makes it work by pairing a downsampling encoder with an upsampling decoder and wiring skip connections between them.
- Vision Transformers (ViT)
Cut the image into patches, treat each patch as a word, run a standard transformer. Don't look back. Implement patch embedding, learned positional embedding, class token, and transformer encoder…
- Diffusion Transformers & Rectified Flow
The U-Net is not the secret of diffusion. Replace it with a transformer, swap the noise schedule for a straight-line flow, and suddenly you have SD3, FLUX, and every 2026 text-to-image model.
- Sequence-to-Sequence Models
Two RNNs pretending to be a translator. The bottleneck they hit is the reason attention exists. Classification maps a variable-length sequence to a single label.
Covered in Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 07: Transformers Deep Dive, Phase 08: Generative AI, Phase 12: Multimodal AI and Phase 19: Capstone Projects.
Related terms
- DecoderA component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and…
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- EmbeddingA learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
- Automatic Speech Recognition (ASR)The task and system pipeline that maps a speech signal to a transcription, often with optional token or segment timing and confidence…
- VAE (Variational Autoencoder)A latent-variable model trained with a reconstruction objective and a regularization term that keeps an approximate posterior close to a…
- Vision Transformer (ViT)A vision architecture that represents an image as a sequence of patch embeddings with position information and processes that sequence…
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.