Scaling and Training Large Models Efficiently episode artwork

EPISODE · Apr 26, 2026 · 14 MIN

Scaling and Training Large Models Efficiently

from Mastering Language Models: From Architecture to Optimization

Topic 2 opens with the question the Transformer made urgent: once you can build big, how should one fixed training budget be split between model size, training tokens, and data quality? Maya and Leo stage the scale-first versus compute-optimal argument in its strongest forms, introduce the smooth-curve predictability of scaling laws, the four interacting knobs of scale, and the two-bills view of training versus inference cost — then map the three deep-dives: Kaplan's scaling laws, Chinchilla's budget correction, and the data-constrained regime where fresh text runs short. Sources: • Scaling Laws for Neural Language Models: https://arxiv.org/pdf/2001.08361 • Training Compute-Optimal Large Language Models: https://arxiv.org/pdf/2203.15556 • Scaling Data-Constrained Language Models: https://arxiv.org/pdf/2305.16264

Episode metadata supplied by the publisher feed · Published Apr 26, 2026

Embed this episode

NOW PLAYING

Scaling and Training Large Models Efficiently

0:00 14:16

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Mastering Language Models: From Architecture to Optimization?

This episode is 14 minutes long.

When was this Mastering Language Models: From Architecture to Optimization episode published?

This episode was published on April 26, 2026.

Can I download this Mastering Language Models: From Architecture to Optimization episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!