Models & inference · Glossary term

What is Autoregressive?

A factorization in which each output token is predicted from the tokens that precede it. During generation, the selected token is appended to the sequence and becomes part of the next prediction's context.

What people say

“The model generates one word at a time.”

What is the common confusion about Autoregressive?

The unit is a token, not necessarily a word, and generation can use decoding methods other than always selecting the highest-probability token.

Learn Autoregressive in the course

Lessons that name Autoregressive in a title or section

  • Visual Autoregressive Modeling (VAR): Next-Scale Prediction

    Diffusion models sample iteratively in time (denoising steps). VAR samples iteratively in scale — it predicts a 1x1 token, then 2x2, then 4x4, up to the final resolution, each scale conditioning on…

    Phase 08: Generative AI

  • Transfusion: Autoregressive Text + Diffusion Image in One Transformer

    Chameleon and Emu3 bet everything on discrete tokens. They work, but the quantization bottleneck is visible — the image quality plateaus below continuous-space diffusion models.

    Phase 12: Multimodal AI

  • Time Series Fundamentals

    Past performance does predict future results -- if you check for stationarity first. Language: Python Decompose a time series into trend, seasonality, and residual components and test for…

    Phase 02: ML Fundamentals

Covered in Phase 02: ML Fundamentals, Phase 08: Generative AI and Phase 12: Multimodal AI.

  • TokenAn integer identifier produced by a model-specific tokenizer from text, bytes, images, audio, or another input representation.
  • TemperatureA decoding parameter that rescales logits before a probability distribution is formed.
  • KV CacheStored key and value tensors from earlier positions in autoregressive generation. Reusing them avoids recomputing attention projections…
  • Decode PhaseThe iterative stage of autoregressive inference that generates new tokens one step at a time after the input prefix has been processed.
  • DecoderA component that maps a representation into an output. In an encoder-decoder transformer, the decoder uses masked self-attention and…
  • Decoding StrategyThe algorithm that converts a model's sequence of next-token scores into selected tokens and a completed output.
  • GPTGenerative Pre-trained Transformer, a family label for generative transformer models pretrained on sequence-prediction objectives and…
  • InferenceExecuting a trained model to produce predictions, scores, embeddings, or generated tokens without performing an ordinary training update…
  • LLM (Large Language Model)A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation.
  • Speculative DecodingAn inference method in which a cheaper draft process proposes several tokens and the target model scores those draft positions in parallel.
  • StreamingDelivering incremental response events before the complete result is ready. A stream may contain token text, structured deltas, tool-call…

More terms in Models & inference

Open the Models & inference list in the glossary

This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.