Prompting & context · Glossary term
What is Token Budget?
An explicit allocation of token capacity across instructions, evidence, history, tool results, reasoning or working space, and output.
Why does Token Budget matter?
Every included token competes for context capacity, latency, and cost. Budgeting forces you to preserve high-value evidence first.
Token Budget in practice
Reserve output capacity, cap retrieved chunks, summarize old tool results into state, and stop or compact before the model limit is reached.
What is the common confusion about Token Budget?
A token budget is a planning constraint. It is not the same as the model's maximum context window.
Learn Token Budget in the course
Start with
- Context Engineering: Windows, Budgets, Memory, and Retrieval
Prompt engineering is a subset. Context engineering is the whole game. A prompt is a string you type. Context is everything that goes into the model's window: system instructions, retrieved…
Lessons that name Token Budget in a title or section
- Any-Resolution Vision: Patch-n'-Pack and NaFlex
Real images are not 224x224 squares. A receipt is 9:16, a chart is 16:9, a medical scan might be 4096x4096, a mobile screenshot is 9:19.5.
- LLaVA-OneVision: Single-Image, Multi-Image, Video in One Model
Before LLaVA-OneVision (Li et al., August 2024) the open-VLM world had separate lineages: LLaVA-1.5 for single images, multi-image models like Mantis and VILA, video models like Video-LLaVA and…
Taught in Phase 11: LLM Engineering.
Also covered in Phase 12: Multimodal AI.
Related terms
- Context WindowThe maximum token capacity available to one model inference under a specific model and API contract.
- Context EngineeringDesigning the full information environment supplied to a model at each step, including instructions, selected files, retrieved evidence,…
- Progressive DisclosureSupplying a person or model with the minimum useful context first, then revealing deeper detail when the task or evidence requires it.
- Cost per Successful TaskTotal system cost divided by the number of tasks that satisfy a defined success criterion, including retries, failed runs, tool use, and…
- Context CompressionReducing the token footprint of source material while attempting to preserve the information required for a later model decision.
- Skill CatalogThe compact model-visible inventory of eligible skills, usually containing routing metadata such as name, description, and an internal…
- Termination ConditionAn explicit rule that ends or pauses an agent run when it succeeds, fails, exhausts a budget, reaches a safe boundary, or requires…
- Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
More terms in Prompting & context
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.