Retrieval & generation · Glossary term

What is Chunking?

Dividing source material into retrievable units before indexing. Chunk boundaries, overlap, metadata, and document structure determine whether retrieval returns enough context without flooding the prompt.

What people say

“Splitting documents into pieces.”

Why does Chunking matter?

The right chunking strategy depends on document shape, query type, embedding model, and evaluation results. There is no universal token size or overlap percentage.

Chunking in practice

Keep headings and code blocks intact, attach source metadata, then measure retrieval quality on real questions before tuning size.

Learn Chunking in the course

Lessons that name Chunking in a title or section

  • Chunking Strategies for RAG

    Chunking configuration influences retrieval quality as much as the choice of embedding model (Vectara NAACL 2025). Get chunking wrong and no amount of reranking saves you.

    Phase 05: NLP: Foundations to Advanced

  • Advanced RAG (Chunking, Reranking, Hybrid Search)

    Basic RAG retrieves the top-k most similar chunks. That works for simple questions. It falls apart for multi-hop reasoning, ambiguous queries, and large corpora.

    Phase 11: LLM Engineering

  • Chunking Strategies, Compared

    Chunking decides what your retriever can ever surface. Get the boundaries wrong and no embedding model, no reranker, no LLM can repair the damage downstream.

    Phase 19: Capstone Projects

  • Build a Voice Assistant Pipeline — The Phase 6 Capstone

    Everything from lessons 01-11, stitched together. Build a voice assistant that listens, reasons, and talks back. In 2026 that is a solved engineering problem, not a research problem — but the…

    Phase 06: Speech & Audio

  • Embeddings & Vector Representations

    Text is discrete. Math is continuous. Every time you ask an LLM to find "similar" documents, compare meanings, or search beyond keywords, you're relying on a bridge between these two worlds.

    Phase 11: LLM Engineering

  • RAG (Retrieval-Augmented Generation)

    Your LLM knows everything up to its training cutoff. It knows nothing about your company's docs, your codebase, or last week's meeting notes.

    Phase 11: LLM Engineering

Covered in Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio, Phase 11: LLM Engineering and Phase 19: Capstone Projects.

  • RAG (Retrieval-Augmented Generation)A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or…
  • RerankerA second-stage model or scoring function that reorders a small candidate set using a richer comparison between the query and each candidate.
  • GroundingConnecting a generated answer or action to evidence, state, or observations that the system can identify and check.
  • Maximum Marginal Relevance (MMR)A selection rule that balances relevance to the query with novelty relative to items already selected.

More terms in Retrieval & generation

Open the Retrieval & generation list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.