Data & representations · Glossary term

What is Embedding?

A learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together

What people say

“A vector that represents meaning.”

What is the common confusion about Embedding?

Similarity depends on the model, training objective, and metric. Distance in one embedding space does not carry over to another.

Why is it called Embedding?

The items are placed, or embedded, in a geometric representation space.

Learn Embedding in the course

Start with

  • Embeddings & Vector Representations

    Text is discrete. Math is continuous. Every time you ask an LLM to find "similar" documents, compare meanings, or search beyond keywords, you're relying on a bridge between these two worlds.

    Phase 11: LLM Engineering

Lessons that name Embedding in a title or section

  • Word Embeddings — Word2Vec from Scratch

    A word is the company it keeps. Train a shallow net on that idea and geometry falls out. TF-IDF knows dog and puppy are different words. It does not know they mean nearly the same thing.

    Phase 05: NLP: Foundations to Advanced

  • GloVe, FastText, and Subword Embeddings

    Word2Vec trained one embedding per word. GloVe factorized the co-occurrence matrix. FastText embedded the pieces. BPE bridged to transformers. Word2Vec left two open questions.

    Phase 05: NLP: Foundations to Advanced

  • Embedding Models — The 2026 Deep Dive

    Word2Vec gave you a vector per word. Modern embedding models give you a vector per passage, cross-lingual, with sparse, dense, and multi-vector views, sized to fit your index.

    Phase 05: NLP: Foundations to Advanced

  • Token and Positional Embeddings

    Ids are integers. The model wants vectors. Two lookup tables sit between them, and the choice of the positional one shapes what the model can learn.

    Phase 19: Capstone Projects

  • Hybrid Retrieval with BM25 and Dense Embeddings

    Lexical and semantic retrieval fail on opposite query distributions. Hybrid retrieval with reciprocal rank fusion does not interpolate, it votes - and the vote wins on every query class.

    Phase 19: Capstone Projects

  • Norms and Distances

    Your distance function defines what "similar" means. Choose wrong and everything downstream breaks. Language: Python Implement L1, L2, cosine, Mahalanobis, Jaccard, and edit distance functions from…

    Phase 01: Math Foundations

  • Vision Transformers (ViT)

    Cut the image into patches, treat each patch as a word, run a standard transformer. Don't look back. Implement patch embedding, learned positional embedding, class token, and transformer encoder…

    Phase 04: Computer Vision

  • Bag of Words, TF-IDF, and Text Representation

    Count first, think later. TF-IDF still beats embeddings on well-defined tasks in 2026. The model needs numbers. You have strings. Every NLP pipeline has to answer the same question.

    Phase 05: NLP: Foundations to Advanced

Taught in Phase 11: LLM Engineering.

Also covered in Phase 01: Math Foundations, Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio, Phase 07: Transformers Deep Dive, Phase 08: Generative AI, Phase 10: LLMs from Scratch, Phase 12: Multimodal AI and Phase 19: Capstone Projects.

  • Cosine SimilarityThe normalized dot product of two vectors. It compares their direction rather than their magnitude and ranges from -1 to 1 for real-valued…
  • Semantic SearchRetrieval that represents a query and candidates in an embedding space and ranks candidates using a vector-similarity function.
  • Vector DatabaseA storage and indexing system that supports nearest-neighbor queries over vector representations, often with metadata filtering,…
  • Audio TokenA discrete identifier produced by an audio codec or tokenizer for a short segment or feature of an audio signal, sometimes across several…
  • Byte Pair Encoding (BPE)A subword-tokenization method that repeatedly merges frequent adjacent units to construct a fixed vocabulary from training text.
  • Contrastive LearningTraining by pulling similar pairs closer and pushing dissimilar pairs apart in embedding space.
  • Dense RetrievalFirst-stage retrieval that embeds queries and candidates into vector representations and ranks candidates by a similarity function.
  • EncoderA component that transforms input into a representation. A transformer encoder commonly uses non-causal self-attention, subject to any…
  • FeatureAn individual measurable property of the data. In classical ML, you engineer features by hand.
  • HNSWAn approximate-nearest-neighbor index that organizes vectors in layered proximity graphs and searches from coarse upper layers toward…
  • Hybrid RetrievalRetrieval that combines signals from different methods, commonly lexical matching and dense-vector similarity, before merging or reranking…
  • Latent SpaceA learned representation space whose coordinates encode factors useful to a model. It may be lower-dimensional than the input, but…
  • ModalityA form of information with its own structure and acquisition process, such as text, image, audio, video, depth, or sensor measurements.
  • Patch EmbeddingA learned projection that converts an image patch into a fixed-width vector used as one element of a transformer input sequence.
  • Semantic CacheA cache that reuses a previous result when a new request is judged sufficiently similar under a chosen representation and threshold.
  • Shared Embedding SpaceA common vector space in which representations from different modalities can be compared with the same similarity function.
  • TokenizationConverting an input representation into the ordered token identifiers a specific model or tokenizer accepts.
  • VocabularyThe finite mapping between token identifiers and the units a tokenizer can emit, including ordinary, byte-level, and special control tokens.

More terms in Data & representations

Open the Data & representations list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.