Models & inference · Glossary term
What is Streaming?
Delivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call arguments, usage metadata, or status events depending on the API.
“Showing output as it is generated.”
Why does Streaming matter?
It improves perceived responsiveness, but it does not reduce the model's actual time to produce a complete answer.
What is the common confusion about Streaming?
Network transport, event shape, and chunk boundaries are provider-specific and are not guaranteed to align with words or tokens.
Learn Streaming in the course
Start with
- Building a Production LLM Application
You have built prompts, embeddings, RAG pipelines, function calling, caching layers, and guardrails. Separately. In isolation. Like practicing guitar scales without ever playing a song.
Lessons that name Streaming in a title or section
- Streaming Speech-to-Speech — Moshi, Hibiki, and Full-Duplex Dialogue
2024-2026 redefined voice AI. Moshi ships a single model that listens and speaks simultaneously at 200 ms latency. Hibiki does speech-to-speech translation chunk-by-chunk.
- MIO and Any-to-Any Streaming Multimodal Models
GPT-4o ships a product most open models cannot replicate: an agent that hears voice, sees video, and speaks back in real time.
- Parallel Tool Calls and Streaming with Tools
Three independent weather lookups serialized is three round trips. Run them in parallel and total time collapses to the slowest single call.
- Speech Recognition (ASR) — CTC, RNN-T, Attention
Speech recognition is audio classification at every timestep, glued together by a sequence model that knows English and silence. CTC, RNN-T, and attention are the three ways to do it.
- Real-Time Audio Processing
Batch pipelines process a file. Real-time pipelines process the next 20 milliseconds before the next 20 arrive. Every conversational AI, broadcast studio, and telephony bot lives and dies by this…
- Build a Voice Assistant Pipeline — The Phase 6 Capstone
Everything from lessons 01-11, stitched together. Build a voice assistant that listens, reasons, and talks back. In 2026 that is a solved engineering problem, not a research problem — but the…
- Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards
You cannot ship what you cannot measure. This lesson names the 2026 metrics for every audio task: ASR (WER, CER, RTFx), TTS (MOS, UTMOS, SECS, WER-on-ASR-round-trip), audio-language (MMAU,…
- Omni Models: Qwen2.5-Omni and the Thinker-Talker Split
GPT-4o's product demo in May 2024 was disruptive not because of the underlying model but because of the product shape — a voice interface where you talk, the model sees what the camera sees, and it…
Taught in Phase 11: LLM Engineering.
Also covered in Phase 06: Speech & Audio, Phase 12: Multimodal AI, Phase 13: Tools & Protocols, Phase 14: Agent Engineering, Phase 17: Infrastructure & Production and Phase 19: Capstone Projects.
Related terms
- Time to First Token (TTFT)The elapsed time from submitting a generation request until the client receives the first output token or content event under a defined…
- AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
- ObservabilityThe ability to understand an AI system's behavior from recorded inputs, outputs, state transitions, tool calls, timings, costs, errors,…
- InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
- Time per Output Token (TPOT)For one request with `N > 1` output tokens, the average post-first-token interval: `(t_N - t_1) / (N - 1)`.
- Tokens per Second (TPS)A throughput measure reporting how many output tokens a serving system produces per unit time under a stated scope and workload.
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.