EPISODE · Jun 10, 2026
Unified Neural Scaling Laws Across Regimes
from AI Post Transformers
This episode explores Unified Neural Scaling Laws, a framework for predicting model performance when parameter count, data volume, training steps, inference compute, and training-recipe choices all change at once. It explains how the paper moves beyond classic smooth power-law curves by introducing broken multivariate scaling laws with regime shifts, including hyperbreaks, bottleneck versus non-bottleneck components, and joint interaction surfaces across training variables. The discussion highlights the paper’s argument that good forecasting must capture both beneficial scaling effects and harmful effects such as overfitting, bad hyperparameter regimes, and data or compute limits, rather than assuming one clean trend forever. A listener would find it interesting because it ties abstract scaling-law math directly to expensive real-world training decisions and to the question of whether pretraining gains actually transfer to downstream benchmarks. Sources: 1. Unified Neural Scaling Laws — Ethan Caballero, Priyank Jaini, David Krueger, Irina Rish, 2026 http://arxiv.org/abs/2605.26248 2. A Constructive Prediction of the Generalization Error Across Scales — Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir Shavit, 2020 https://scholar.google.com/scholar?q=A+Constructive+Prediction+of+the+Generalization+Error+Across+Scales 3. Scaling Laws for Neural Language Models — Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, et al., 2020 https://scholar.google.com/scholar?q=Scaling+Laws+for+Neural+Language+Models 4. Training Compute-Optimal Large Language Models — Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al., 2022 https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models 5. Unified Neural Scaling Laws — Ethan Caballero, Priyank Jaini, David Krueger, Irina Rish, 2026 https://scholar.google.com/scholar?q=Unified+Neural+Scaling+Laws 6. Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour — Priya Goyal, Piotr Dollar, Ross Girshick, Kaiming He, et al., 2017 https://scholar.google.com/scholar?q=Accurate%2C+Large+Minibatch+SGD%3A+Training+ImageNet+in+1+Hour 7. An Empirical Model of Large-Batch Training — Sam McCandlish, Jared Kaplan, Dario Amodei, OpenAI Dota Team, 2018 https://scholar.google.com/scholar?q=An+Empirical+Model+of+Large-Batch+Training 8. Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer — Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, et al., 2022 https://scholar.google.com/scholar?q=Tensor+Programs+V%3A+Tuning+Large+Neural+Networks+via+Zero-Shot+Hyperparameter+Transfer 9. Resolving Discrepancies in Compute-Optimal Scaling of Language Models — Tomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt, Yair Carmon, 2024 https://scholar.google.com/scholar?q=Resolving+Discrepancies+in+Compute-Optimal+Scaling+of+Language+Models 10. Reconciling modern machine learning practice and the bias-variance trade-off — Mikhail Belkin, Daniel Hsu, Siyuan Ma, Soumik Mandal, 2019 https://scholar.google.com/scholar?q=Reconciling+modern+machine+learning+practice+and+the+bias-variance+trade-off 11. Deep Double Descent: Where Bigger Models and More Data Hurt — Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, Ilya Sutskever, 2020 https://scholar.google.com/scholar?q=Deep+Double+Descent%3A+Where+Bigger+Models+and+More+Data+Hurt 12. Broken Neural Scaling Laws — Ethan Caballero, Kshitij Gupta, Irina Rish, David Krueger, 2023 https://scholar.google.com/scholar?q=Broken+Neural+Scaling+Laws 13. Scaling Data-Constrained Language Models — Niklas Muennighoff, Alexander M. Rush, Boaz Barak, et al., 2023 https://scholar.google.com/scholar?q=Scaling+Data-Constrained+Language+Models 14. Observational Scaling Laws and the Predictability of Language Model Performance — Yangjun Ruan, Chris J. Maddison, Tatsunori Hashimoto, 2024 https://scholar.google.com/scholar?q=Observational+Scaling+Laws+and+the+Predictability+of+Language+Model+Performance 15. Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic — Sachin Goyal et al., 2024 https://scholar.google.com/scholar?q=Scaling+Laws+for+Data+Filtering+--+Data+Curation+cannot+be+Compute+Agnostic 16. Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies — Zhengyu Chen et al., 2025 https://scholar.google.com/scholar?q=Revisiting+Scaling+Laws+for+Language+Models%3A+The+Role+of+Data+Quality+and+Training+Strategies 17. Scaling Laws for Optimal Data Mixtures — Mustafa Shukor et al., 2025 https://scholar.google.com/scholar?q=Scaling+Laws+for+Optimal+Data+Mixtures 18. Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient — Jan Ludziejewski et al., 2025 https://scholar.google.com/scholar?q=Joint+MoE+Scaling+Laws%3A+Mixture+of+Experts+Can+Be+Memory+Efficient 19. Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning — Wenkai Yang et al., 2025 https://scholar.google.com/scholar?q=Towards+Thinking-Optimal+Scaling+of+Test-Time+Compute+for+LLM+Reasoning 20. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping et al., 2025 https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach 21. Towards Robust Scaling Laws for Optimizers — Alexandra Volkova et al., 2026 https://scholar.google.com/scholar?q=Towards+Robust+Scaling+Laws+for+Optimizers 22. Neural Scaling Laws Rooted in the Data Distribution — Ari Brill, 2024 https://scholar.google.com/scholar?q=Neural+Scaling+Laws+Rooted+in+the+Data+Distribution 23. The Quantization Model of Neural Scaling — Eric J. Michaud et al., 2023 https://scholar.google.com/scholar?q=The+Quantization+Model+of+Neural+Scaling 24. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3 25. AI Post Transformers: TMAS: Scaling Test-Time Compute with Multi-Agent Synergy — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-14-tmas-scaling-test-time-compute-with-mult-3abe7a.mp3 26. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3 Interactive Visualization: Unified Neural Scaling Laws Across Regimes
Embed this episode
NOW PLAYING
Unified Neural Scaling Laws Across Regimes
No transcript for this episode yet
Similar Episodes
No similar episodes found.