Scaling Laws
The 2020 Kaplan paper said: bigger model, lower loss. The 2022 Hoffmann paper said: you were under-training. Compute goes into two buckets — parameters and tokens — and the split is not obvious. When you have C FLOPs of training compute and want the best model, you face two knobs: How many parameters (N)? Bigger model, higher capacity. How many training tokens (D)? More data, better use of capacity. FLOPs scale approximately as 6 × N × D. You can push N up and D down, or D up and N down. Which is better? Before 2022, the answer was "push N hard." GPT-3 (2020) was 175B parameters trained on 300B tokens. A ratio of about 1.7 tokens per parameter. The Kaplan scaling laws backed this up. Hoffmann et al. (2022), training a small family of models called Chinchilla, found something different: optimal ratio is closer to 20 tokens per parameter. GPT-3 was 10× undertrained. Chinchilla (70B params, 1.4T tokens) beat GPT-3 (175B, 300B tokens) on every benchmark at 2.5× less inference cost. 2026 is Chinchilla's world — with one important twist. Llama 3 8B was trained on 15 trillion tokens, a ratio of 1,875 tokens per parameter. Ninety-four times past Chinchilla-optimal. Inference cost matters more than training cost for models that will be used at scale, so over-training (past…
Scaling Laws: The 2020 Kaplan paper said: bigger model, lower loss. The 2022 Hoffmann paper said: you were under-training. Compute goes into two buckets —…
This free lesson is part of the AI Engineering from Scratch curriculum. Read the full explanation, run the lesson code, and verify the result in the interactive reader or from the repository source.
Browse the complete course catalog or open this lesson on GitHub.