Models & inference · Glossary term

What is Streaming?

Delivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call arguments, usage metadata, or status events depending on the API.

What people say

“Showing output as it is generated.”

Why does Streaming matter?

It improves perceived responsiveness, but it does not reduce the model's actual time to produce a complete answer.

What is the common confusion about Streaming?

Network transport, event shape, and chunk boundaries are provider-specific and are not guaranteed to align with words or tokens.

Learn Streaming in the course

Start with

  • Building a Production LLM Application

    You have built prompts, embeddings, RAG pipelines, function calling, caching layers, and guardrails. Separately. In isolation. Like practicing guitar scales without ever playing a song.

    Phase 11: LLM Engineering

Lessons that name Streaming in a title or section

Taught in Phase 11: LLM Engineering.

Also covered in Phase 06: Speech & Audio, Phase 12: Multimodal AI, Phase 13: Tools & Protocols, Phase 14: Agent Engineering, Phase 17: Infrastructure & Production and Phase 19: Capstone Projects.

  • Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
  • AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
  • ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
  • InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
  • Time per Output Token (TPOT)For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`.
  • Tokens per Second (TPS)A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.