Scaling Laws for Neural Language Models episode artwork

EPISODE · Apr 26, 2026 · 14 MIN

Scaling Laws for Neural Language Models

from Mastering Language Models: From Architecture to Optimization

Deep dive into Kaplan et al.'s Scaling Laws for Neural Language Models (2020), the paper that made giant training runs forecastable. Maya and Leo walk the three landmarks: the ruler — loss falls along smooth power laws in parameters, data, and compute, so cheap pilot runs predict frontier runs; the early exit — larger models learn more per token, so a fixed budget should buy a huge model trained on modest data and stopped before convergence; and the edge of the map — loss is a proxy, curves are fitted to a measured range, and averages can hide brittle rare-task behavior. They stage the real argument between curve-trusting planners and loss-as-proxy skeptics, and set up Chinchilla's revision next episode. Sources: • Scaling Laws for Neural Language Models: https://arxiv.org/pdf/2001.08361 • Training Compute-Optimal Large Language Models: https://arxiv.org/pdf/2203.15556

Episode metadata supplied by the publisher feed · Published Apr 26, 2026

Embed this episode

NOW PLAYING

Scaling Laws for Neural Language Models

0:00 14:04

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Mastering Language Models: From Architecture to Optimization?

This episode is 14 minutes long.

When was this Mastering Language Models: From Architecture to Optimization episode published?

This episode was published on April 26, 2026.

Can I download this Mastering Language Models: From Architecture to Optimization episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!