Models & inference · Glossary term

What is Inference?

Executing a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update to its parameters.

What people say

“Running a trained model.”

What is the common confusion about Inference?

An application can update caches, conversation state, or external memory during inference even though model weights stay unchanged.

Learn Inference in the course

Lessons that name Inference in a title or section

  • Natural Language Inference — Textual Entailment

    "t entails h" means a human reading t would conclude h is true. NLI is the task of predicting entailment / contradiction / neutral. Boring on the surface, load-bearing in production.

    Phase 05: NLP: Foundations to Advanced

  • KV Cache, Flash Attention & Inference Optimization

    Training is parallel and FLOP-bound. Inference is serial and memory-bound. Different bottleneck, different tricks. A naive autoregressive decoder does O(N²) work to generate N tokens: at each step…

    Phase 07: Transformers Deep Dive

  • Inference Optimization

    Two phases define LLM inference. Prefill processes your prompt in parallel -- compute-bound. Decode generates tokens one at a time -- memory-bound. Every optimization targets one or both.

    Phase 10: LLMs from Scratch

  • Async and Hogwild! Inference

    Speculative decoding (Phase 10 · 15) parallelizes tokens within one sequence. Multi-agent frameworks parallelize across whole sequences but force explicit coordination (voting, sub-task splitting).

    Phase 10: LLMs from Scratch

  • Inference Platform Economics — Fireworks, Together, Baseten, Modal, Replicate, Anyscale

    The 2026 inference market is no longer GPU time rental. It bifurcates into custom silicon (Groq, Cerebras, SambaNova), GPU platforms (Baseten, Together, Fireworks, Modal), and API-first marketplaces…

    Phase 17: Infrastructure & Production

  • Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell

    Hardware-specialized inference compilation trades portability for throughput, and TensorRT-LLM — NVIDIA-only, tuned for Blackwell — is the clearest example of the trade paying off.

    Phase 17: Infrastructure & Production

  • Inference Metrics — TTFT, TPOT, ITL, Goodput, P99

    Four metrics decide whether an inference deployment is working. TTFT is prefill plus queue plus network. TPOT (equivalently ITL) is the memory-bound decode cost per token.

    Phase 17: Infrastructure & Production

  • Edge Inference — Apple Neural Engine, Qualcomm Hexagon, WebGPU/WebLLM, Jetson

    The core edge constraint is memory bandwidth, not compute. Mobile DRAM sits at 50-90 GB/s; datacenter HBM3 clears 2-3 TB/s — a 30-50x gap. Decode is memory-bound so the gap is decisive.

    Phase 17: Infrastructure & Production

Covered in Phase 04: Computer Vision, Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio, Phase 07: Transformers Deep Dive, Phase 08: Generative AI, Phase 10: LLMs from Scratch, Phase 17: Infrastructure & Production and Phase 19: Capstone Projects.

  • AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • Diffusion ModelA generative model trained around a progressive noising process and a learned reverse process.
  • Model ServingThe runtime and API layer that loads versioned model artifacts, accepts inference requests, schedules execution, manages resources, and…

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.