Compute-Optimal Scaling Laws for Language Models Revisited episode artwork

EPISODE · Apr 18, 2025 · 17 MIN

Compute-Optimal Scaling Laws for Language Models Revisited

from Best AI papers explained · host Enoch H. Kang

This paper investigates discrepancies in scaling laws for compute-optimal language models, particularly between Kaplan et al. and Hoffmann et al. The authors reproduce the Kaplan et al. law and identify key factors causing the divergence: the computational cost of the last layer, the length of the learning rate warmup, and the importance of scale-dependent optimizer tuning. After correcting for these elements, the study achieves strong agreement with the Hoffmann et al. scaling law, notably demonstrating that specific learning rate decay schedules are not essential. Additionally, the research derives scaling laws for optimal learning rates and batch sizes, highlighting the significance of tuning the AdamW $\beta_2$ parameter at smaller batch sizes.

Episode metadata supplied by the publisher feed · Published Apr 18, 2025

Embed this episode

NOW PLAYING

Compute-Optimal Scaling Laws for Language Models Revisited

0:00 17:14

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Best AI papers explained?

This episode is 17 minutes long.

When was this Best AI papers explained episode published?

This episode was published on April 18, 2025.

Can I download this Best AI papers explained episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!