Models & inference · Glossary term
What is Inference?
Executing a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update to its parameters.
“Running a trained model.”
What is the common confusion about Inference?
An application can update caches, conversation state, or external memory during inference even though model weights stay unchanged.
Learn Inference in the course
Lessons that name Inference in a title or section
- Natural Language Inference — Textual Entailment
"t entails h" means a human reading t would conclude h is true. NLI is the task of predicting entailment / contradiction / neutral. Boring on the surface, load-bearing in production.
- KV Cache, Flash Attention & Inference Optimization
Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different tricks. A naive autoregressive decoder does O(N²) work to generate N tokens: at each step…
- Inference Optimization
Two phases define LLM inference. Prefill processes your prompt in parallel -- compute-bound. Decode generates tokens one at a time -- memory-bound. Every optimization targets one or both.
- Async and Hogwild! Inference
Speculative decoding (Phase 10 · 15) parallelizes tokens within one sequence. Multi-agent frameworks parallelize across whole sequences but force explicit coordination (voting, sub-task splitting).
- Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscale
The 2026 inference market is no longer GPU time rental. It bifurcates into custom silicon (Groq, Cerebras, SambaNova), GPU platforms (Baseten, Together, Fireworks, Modal), and API-first marketplaces…
- Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell
Hardware-specialized inference compilation trades portability for throughput, and TensorRT-LLM — NVIDIA-only, tuned for Blackwell — is the clearest example of the trade paying off.
- Inference Metrics — TTFT, TPOT, ITL, Goodput, P99
Four metrics decide whether an inference deployment is working. TTFT is prefill plus queue plus network. TPOT (equivalently ITL) is the memory-bound decode cost per token.
- Edge Inference — Apple Neural Engine, Qualcomm Hexagon, WebGPU/WebLLM, Jetson
The core edge constraint is memory bandwidth, not compute. Mobile DRAM sits at 50-90 GB/s; datacenter HBM3 clears 2-3 TB/s — a 30-50x gap. Decode is memory-bound so the gap is decisive.
Covered in Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio, Phase 07: Transformers Deep Dive, Phase 08: Generative AI, Phase 10: LLMs from Scratch, Phase 17: Infrastructure & Production and Phase 19: Capstone Projects.
Related terms
- AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
- StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
- KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
- Diffusion ModelA generative model trained around a progressive noising process and a learned reverse process.
- Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.