Prompting & context · Glossary term
What is Context Window?
The maximum token capacity available to one model inference under a specific model and API contract. The capacity may include system instructions, messages, retrieved content, tool exchanges, and generated output, with provider-specific accounting and output limits.
“How much the model remembers.”
Why does Context Window matter?
Conversation history is only available when the application sends or reconstructs it. A large window does not guarantee that every included detail will be used reliably.
What is the common confusion about Context Window?
Context is temporary input to an inference. Durable memory is stored outside the model and selected back into later context.
Learn Context Window in the course
Start with
- Context Engineering: Windows, Budgets, Memory, and Retrieval
Prompt engineering is a subset. Context engineering is the whole game. A prompt is a string you type. Context is everything that goes into the model's window: system instructions, retrieved…
Lessons that name Context Window in a title or section
- Prompt Engineering: Techniques & Patterns
Most people write prompts like they are texting a friend. Then they wonder why a 200-billion parameter model gives mediocre answers. Prompt engineering is not about tricks.
Taught in Phase 11: LLM Engineering.
Related terms
- Token BudgetAn explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
- Context EngineeringDesigning the full information environment supplied to a model at each step, including instructions, selected files, retrieved evidence,…
- Prompt CacheReuse of provider-side or application-side computation for an identical or eligible prompt prefix so repeated inference avoids some…
- Agent StateThe explicit data an agent carries across steps, such as the current objective, completed actions, tool results, open questions, budgets,…
- Few-ShotIn-context learning that includes a small set of demonstrations before the target input so the model can infer the desired task, format,…
- In-Context LearningA model adapting its behavior from instructions, examples, or patterns supplied in the current input without an ordinary parameter update.
- Lost in the MiddleA long-context failure pattern in which model performance changes with evidence position and can degrade when relevant information sits…
- Paged KV CacheA KV-cache memory manager that stores attention state in fixed-size blocks and maps logical sequence positions to physical blocks instead…
- Self-AttentionAttention in which queries, keys, and values are derived from the same sequence representation.
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
More terms in Prompting & context
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.