Models & inference · Glossary term
What is Autoregressive?
A factorization in which each output token is predicted from the tokens that precede it. During generation, the selected token is appended to the sequence and becomes part of the next prediction's context.
“The model generates one word at a time.”
What is the common confusion about Autoregressive?
The unit is a token, not necessarily a word, and generation can use decoding methods other than always selecting the highest-probability token.
Learn Autoregressive in the course
Lessons that name Autoregressive in a title or section
- Visual Autoregressive Modeling (VAR): Next-Scale Prediction
Diffusion models sample iteratively in time (denoising steps). VAR samples iteratively in scale — it predicts a 1x1 token, then 2x2, then 4x4, up to the final resolution, each scale conditioning on…
- Transfusion: Autoregressive Text + Diffusion Image in One Transformer
Chameleon and Emu3 bet everything on discrete tokens. They work, but the quantization bottleneck is visible — the image quality plateaus below continuous-space diffusion models.
- Time Series Fundamentals
Past performance does predict future results -- if you check for stationarity first. Language: Python Decompose a time series into trend, seasonality, and residual components and test for…
Covered in Phase 02: ML Fundamentals, Phase 08: Generative AI and Phase 12: Multimodal AI.
Related terms
- TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
- TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
- KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
- Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
- DecoderA component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and…
- Decoding StrategyThe algorithm that converts a model's sequence of next-token scores into selected tokens and a completed output.
- GPTGenerative Pre-trained Transformer, a family label for generative transformer models pretrained on sequence-prediction objectives and…
- InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
- LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
- Speculative DecodingAn inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel.
- StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…
More terms in Models & inference
- Attention
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.