Scaling laws and compute-optimal training

Hoffmann et al. (Chinchilla, 2022) showed that most models are undertrained — trained on too few tokens relative to their size. The compute optimal ratio is approximately 20 tokens per parameter (not the 1 2 tokens per parameter common in early LLMs).