Retrieval & generation · Glossary term

What is RAG (Retrieval-Augmented Generation)?

A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or acts. Retrieval can use lexical, vector, structured, or hybrid methods.

What people say

“A model answering with retrieved knowledge.”

Why does RAG (Retrieval-Augmented Generation) matter?

RAG can make current or private evidence available without encoding it into model weights, but retrieval and grounding must be evaluated separately.

Why is it called RAG (Retrieval-Augmented Generation)?

Retrieval finds evidence, augmentation adds selected evidence to context, and generation produces the response.

Learn RAG (Retrieval-Augmented Generation) in the course

Start with

  • RAG (Retrieval-Augmented Generation)

    Your LLM knows everything up to its training cutoff. It knows nothing about your company's docs, your codebase, or last week's meeting notes.

    Phase 11: LLM Engineering

Lessons that name RAG (Retrieval-Augmented Generation) in a title or section

  • Chunking Strategies for RAG

    Chunking configuration influences retrieval quality as much as the choice of embedding model (Vectara NAACL 2025). Get chunking wrong and no amount of reranking saves you.

    Phase 05: NLP: Foundations to Advanced

  • Advanced RAG (Chunking, Reranking, Hybrid Search)

    Basic RAG retrieves the top-k most similar chunks. That works for simple questions. It falls apart for multi-hop reasoning, ambiguous queries, and large corpora.

    Phase 11: LLM Engineering

  • ColPali and Vision-Native Document RAG

    Traditional RAG parses PDFs into text, splits into chunks, embeds chunks, stores vectors. Every step loses signal: OCR drops chart data, chunking breaks table rows, text embeddings ignore figures.

    Phase 12: Multimodal AI

  • Multimodal RAG and Cross-Modal Retrieval

    Vision-native document RAG is one slice. Production multimodal RAG goes wider — retrieving across text, images, audio, and video for workflows like trip planning ("find me a quiet vegan brunch with…

    Phase 12: Multimodal AI

  • Capstone 02 — RAG over Codebase (Cross-Repo Semantic Search)

    Every serious engineering org in 2026 runs an internal code search that understands meaning, not just strings. Sourcegraph Amp, Cursor's codebase answers, Augment's enterprise graph, Aider's…

    Phase 19: Capstone Projects

  • Capstone 08 — Production RAG Chatbot for a Regulated Vertical

    Harvey, Glean, Mendable, and LlamaCloud all run the same production shape in 2026. Ingest with docling or Unstructured and ColPali for visuals. Hybrid search. Re-rank with bge-reranker-v2-gemma.

    Phase 19: Capstone Projects

  • RAG Evaluation: Precision, Recall, MRR, nDCG, Faithfulness, Answer Relevance

    If you cannot grade your retrieval and your answer at the same time, you cannot ship the system. The two are not the same metric and the same prompt fails on different axes.

    Phase 19: Capstone Projects

  • End-to-End RAG System

    Six lessons of components. One pipeline. One eval loop. One self-terminating demo. This is the system you ship. Compose the chunker, hybrid retriever, query rewriter, cross-encoder reranker, and…

    Phase 19: Capstone Projects

Taught in Phase 11: LLM Engineering.

Also covered in Phase 05: NLP: Foundations to Advanced, Phase 10: LLMs from Scratch, Phase 12: Multimodal AI and Phase 19: Capstone Projects.

  • GroundingConnecting a generated answer or action to evidence, state, or observations that the system can identify and check.
  • Hybrid RetrievalRetrieval that combines signals from different methods, commonly lexical matching and dense-vector similarity, before merging or reranking…
  • RerankerA second-stage model or scoring function that reorders a small candidate set using a richer comparison between the query and each candidate.
  • HallucinationGenerated content that is false, unsupported by the available evidence, or inconsistent with the task's source of truth.
  • BM25A lexical ranking function that scores a document from query-term matches while accounting for term rarity, repeated occurrences, and…
  • ChunkingDividing source material into retrievable units before indexing. Chunk boundaries, overlap, metadata, and document structure determine…
  • Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
  • Maximum Marginal Relevance (MMR)A selection rule that balances relevance to the query with novelty relative to items already selected.

Sources

More terms in Retrieval & generation

Open the Retrieval & generation list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.