Models & inference · Glossary term
What is CUDA?
NVIDIA's platform and programming model for general-purpose computation on compatible GPUs. Deep-learning frameworks use CUDA libraries and kernels to execute many tensor operations in parallel.
“GPU programming.”
What is the common confusion about CUDA?
GPU acceleration is not synonymous with CUDA; other hardware and software stacks exist.
Learn CUDA in the course
Lessons that name CUDA in a title or section
- Python Environments
Dependency hell is real. Virtual environments are the cure. Create isolated virtual environments using uv, venv, or conda.
Covered in Phase 00: Setup & Tooling.
Related terms
- TensorA typed array with a shape, data type, and device placement that frameworks use to represent inputs, parameters, activations, and gradients.
- Mixed PrecisionA numerical strategy that uses different data types for different operations, often lower precision for many matrix operations and higher…
- JAXA Python library for transforming numerical functions with automatic differentiation, compilation, vectorization, and parallel execution…
More terms in Models & inference
- Attention
- Autoregressive
- CNN (Convolutional Neural Network)
- Decoder
- Decoding Strategy
- Diffusion Model
- Encoder
- GAN (Generative Adversarial Network)
- GPT
- Inductive Bias
- Inference
- KV Cache
- LLM (Large Language Model)
- Logits
- MoE (Mixture of Experts)
- Nucleus Sampling (Top-p)
- Parameter
- Perplexity
- Quantization
- Self-Attention
- Speculative Decoding
- Stop Sequence
- Streaming
- Temperature
- Time to First Token (TTFT)
- Top-k Sampling
- Transformer
- VAE (Variational Autoencoder)
This entry comes from glossary/terms.md on GitHub. Browse all 250 glossary terms.