Models & inference · Glossary term
What is LLM (Large Language Model)?
A language model with enough capacity and broad training to perform many language tasks through prompting or adaptation. Most current LLMs use transformer architectures and sequence-prediction objectives, but size thresholds, data sources, and training recipes vary.
“The brain of an AI application.”
What is the common confusion about LLM (Large Language Model)?
An LLM is a model component. Tools, retrieval, state, policies, and product logic live in the surrounding system.
Learn LLM (Large Language Model) in the course
Lessons that name LLM (Large Language Model) in a title or section
- Chatbots — Rule-Based to Neural to LLM Agents
ELIZA replied with pattern matches. DialogFlow mapped intents. GPT answered from weights. Claude runs tools and verifies. Each era solved the previous one's worst failure.
- LLM Evaluation — RAGAS, DeepEval, G-Eval
Exact-match and F1 miss semantic equivalence. Human review does not scale. LLM-as-judge is the production answer — with enough calibration to trust the number.
- Building a Complete LLM Pipeline
Everything from Lessons 01 to 12 is one stage of one pipeline. This lesson is the scaffold that turns those stages into a single end-to-end run: tokenize, pre-train, scale, SFT, align, evaluate,…
- Evaluation & Testing LLM Applications
You would never deploy a web app without tests. You would never ship a database migration without a rollback plan. But right now, most teams ship LLM applications by reading 10 outputs and saying…
- Building a Production LLM Application
You have built prompts, embeddings, RAG pipelines, function calling, caching layers, and guardrails. Separately. In isolation. Like practicing guitar scales without ever playing a song.
- LLM Routing Layer — LiteLLM, OpenRouter, Portkey
Provider lock-in is expensive. Different tool-calling workloads suit different models. Routing gateways give one API surface, retries, failover, cost tracking, and guardrails.
- Swarm Optimization for LLMs (PSO, ACO)
Bio-inspired optimization is making an LLM comeback. LMPSO (arXiv:2504.09247) uses PSO where each particle's velocity is a prompt and the LLM generates the next candidate; works well on…
- Managed LLM Platforms — Bedrock, Vertex AI, Azure OpenAI
Three hyperscalers, three distinct strategies. AWS Bedrock is a model marketplace — Claude, Llama, Titan, Stability, Cohere behind one API.
Covered in Phase 01: Math Foundations, Phase 05: NLP: Foundations to Advanced, Phase 06: Speech & Audio, Phase 10: LLMs from Scratch, Phase 11: LLM Engineering, Phase 12: Multimodal AI, Phase 13: Tools & Protocols, Phase 14: Agent Engineering, Phase 15: Autonomous Systems, Phase 16: Multi-Agent & Swarms, Phase 17: Infrastructure & Production, Phase 18: Ethics, Safety & Alignment and Phase 19: Capstone Projects.
Related terms
- TransformerA neural-network architecture built from attention, position information, feed-forward sublayers, residual connections, and normalization.
- AutoregressiveA factorization in which each output token is predicted from the tokens that precede it.
- Agent HarnessThe runtime around a model that assembles context, exposes tools, manages state, enforces limits, records traces, and decides when the…
- GPTGenerative Pre-trained Transformer, a family label for generative transformer models pretrained on sequence-prediction objectives and…
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- CUDA
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.