EPISODE · Aug 20, 2026
Scalable Second-Order Optimization: Shampoo at Scale
from AI Post Transformers
This episode explores the paper "Scalable Second Order Optimization for Deep Learning" and its introduction of Distributed Shampoo, a Kronecker-factored second-order optimizer that cuts training steps in half compared to a well-tuned Adam baseline on WMT'14 English-to-French translation, with further gains shown on BERT, Criteo click-through-rate modeling, and ResNet-50. The discussion traces the lineage from Newton's method and full-matrix AdaGrad's prohibitive O(N²)/O(N³) costs through Shampoo's Kronecker-product approximation, which replaces one massive preconditioner with smaller per-dimension matrices to make second-order optimization tractable at scale. Rival factored approaches, K-FAC and K-BFGS, are positioned as points of comparison throughout the paper. The conversation is aimed at listeners who use Adam daily but have never unpacked the mechanical distinction between first-order and second-order optimization, or the meaning of "preconditioning" itself. It's a compelling listen because second-order methods have long been dismissed as theoretically superior but practically unscalable, and this paper demonstrates real wall-clock wins across four distinct production-scale workloads rather than a single cherry-picked benchmark. Sources: 1. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2020 http://arxiv.org/abs/2002.09018 2. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018 https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization 3. Optimizing Neural Networks with Kronecker-factored Approximate Curvature (K-FAC) — James Martens, Roger Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature+%28K-FAC%29 4. Adam: A Method for Stochastic Optimization — Diederik P. Kingma, Jimmy Ba, 2014 https://scholar.google.com/scholar?q=Adam%3A+A+Method+for+Stochastic+Optimization 5. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization (AdaGrad) — John Duchi, Elad Hazan, Yoram Singer, 2011 https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization+%28AdaGrad%29 6. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — J. Martens, R. Grosse, 2015 https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature 7. Distributed Second-Order Optimization using Kronecker-Factored Approximations — J. Ba, J. Martens, R. Grosse, 2017 https://scholar.google.com/scholar?q=Distributed+Second-Order+Optimization+using+Kronecker-Factored+Approximations 8. Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks — K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, R. Yokota, S. Matsuoka, 2019 https://scholar.google.com/scholar?q=Large-Scale+Distributed+Second-Order+Optimization+Using+Kronecker-Factored+Approximate+Curvature+for+Deep+Convolutional+Neural+Networks 9. Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes — Y. You, J. Li, S. Reddi, et al. (LAMB), 2019 https://scholar.google.com/scholar?q=Large+Batch+Optimization+for+Deep+Learning%3A+Training+BERT+in+76+Minutes 10. A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes — Z. Nado, J. Gilmer, C. Shallue, R. Anil, G. Dahl, 2021 https://scholar.google.com/scholar?q=A+Large+Batch+Optimizer+Reality+Check%3A+Traditional%2C+Generic+Optimizers+Suffice+Across+Batch+Sizes 11. Limitations of the Empirical Fisher Approximation for Natural Gradient Descent — F. Kunstner, P. Hennig, L. Balles, 2019 https://scholar.google.com/scholar?q=Limitations+of+the+Empirical+Fisher+Approximation+for+Natural+Gradient+Descent Interactive Visualization: Scalable Second-Order Optimization: Shampoo at Scale
Embed this episode
NOW PLAYING
Scalable Second-Order Optimization: Shampoo at Scale
No transcript for this episode yet
Similar Episodes
No similar episodes found.