Retrieval & generation · Glossary term
What is RAG (Retrieval-Augmented Generation)?
A system pattern that retrieves evidence relevant to a request and supplies selected content to a generative model before it answers or acts. Retrieval can use lexical, vector, structured, or hybrid methods.
“A model answering with retrieved knowledge.”
Why does RAG (Retrieval-Augmented Generation) matter?
RAG can make current or private evidence available without encoding it into model weights, but retrieval and grounding must be evaluated separately.
Why is it called RAG (Retrieval-Augmented Generation)?
Retrieval finds evidence, augmentation adds selected evidence to context, and generation produces the response.
Learn RAG (Retrieval-Augmented Generation) in the course
Start with
- RAG (Retrieval-Augmented Generation)
Your LLM knows everything up to its training cutoff. It knows nothing about your company's docs, your codebase, or last week's meeting notes.
Lessons that name RAG (Retrieval-Augmented Generation) in a title or section
- Chunking Strategies for RAG
Chunking configuration influences retrieval quality as much as the choice of embedding model (Vectara NAACL 2025). Get chunking wrong and no amount of reranking saves you.
- Advanced RAG (Chunking, Reranking, Hybrid Search)
Basic RAG retrieves the top-k most similar chunks. That works for simple questions. It falls apart for multi-hop reasoning, ambiguous queries, and large corpora.
- ColPali and Vision-Native Document RAG
Traditional RAG parses PDFs into text, splits into chunks, embeds chunks, stores vectors. Every step loses signal: OCR drops chart data, chunking breaks table rows, text embeddings ignore figures.
- Multimodal RAG and Cross-Modal Retrieval
Vision-native document RAG is one slice. Production multimodal RAG goes wider — retrieving across text, images, audio, and video for workflows like trip planning ("find me a quiet vegan brunch with…
- Capstone 02 — RAG over Codebase (Cross-Repo Semantic Search)
Every serious engineering org in 2026 runs an internal code search that understands meaning, not just strings. Sourcegraph Amp, Cursor's codebase answers, Augment's enterprise graph, Aider's…
- Capstone 08 — Production RAG Chatbot for a Regulated Vertical
Harvey, Glean, Mendable, and LlamaCloud all run the same production shape in 2026. Ingest with docling or Unstructured and ColPali for visuals. Hybrid search. Re-rank with bge-reranker-v2-gemma.
- RAG Evaluation: Precision, Recall, MRR, nDCG, Faithfulness, Answer Relevance
If you cannot grade your retrieval and your answer at the same time, you cannot ship the system. The two are not the same metric and the same prompt fails on different axes.
- End-to-End RAG System
Six lessons of components. One pipeline. One eval loop. One self-terminating demo. This is the system you ship. Compose the chunker, hybrid retriever, query rewriter, cross-encoder reranker, and…
Taught in Phase 11: LLM Engineering.
Also covered in Phase 05: NLP: Foundations to Advanced, Phase 10: LLMs from Scratch, Phase 12: Multimodal AI and Phase 19: Capstone Projects.
Related terms
- GroundingConnecting a generated answer or action to evidence, state, or observations that the system can identify and check.
- Hybrid RetrievalRetrieval that combines signals from different methods, commonly lexical matching and dense-vector similarity, before merging or reranking…
- RerankerA second-stage model or scoring function that reorders a small candidate set using a richer comparison between the query and each candidate.
- HallucinationGenerated content that is false, unsupported by the available evidence, or inconsistent with the task's source of truth.
- BM25A lexical ranking function that scores a document from query-term matches while accounting for term rarity, repeated occurrences, and…
- ChunkingDividing source material into retrievable units before indexing. Chunk boundaries, overlap, metadata, and document structure determine…
- Fine-tuningContinuing training from pretrained parameters on a narrower dataset or objective. Depending on the method, you may update all parameters,…
- Maximum Marginal Relevance (MMR)A selection rule that balances relevance to the query with novelty relative to items already selected.
Sources
More terms in Retrieval & generation
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.