SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam episode artwork

EPISODE · Aug 21, 2026

SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam

from AI Post Transformers

This episode explores SOAP, a new optimizer from a Harvard/Kempner Institute team that fuses Shampoo's second-order preconditioning with Adam's update mechanics. The hosts trace the lineage from Adagrad's mathematically ideal but computationally infeasible full preconditioner matrix, through Adam's cheap diagonal approximation, to Shampoo's middle-ground Kronecker-product approach using two smaller per-dimension preconditioners. The core theoretical result discussed is a proof that Shampoo run with the one-half power is mathematically equivalent to running Adafactor inside the eigenbasis Shampoo's own preconditioner defines — which motivates simply swapping in full Adam within that same rotated basis, adding just one new hyperparameter (preconditioning frequency) over standard AdamW. The discussion highlights the paper's striking efficiency claims — over 40% fewer training iterations and 35% less wall-clock time versus AdamW, and roughly 20% better than Shampoo itself — while noting these numbers deserve scrutiny given real-world context like Shampoo's AlgoPerf benchmark win and its use in training Gemini 1.5 Flash. Listeners interested in the mechanics behind large-scale training efficiency will get a clear breakdown of why optimizer choice translates directly into cluster-scale compute costs and calendar time. Sources: 1. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Mujin Kwun, Itai Shapira, David Brandfonbrener, Lucas Janson, Sham Kakade, 2024 http://arxiv.org/abs/2409.11321 2. Muon: An optimizer for hidden layers in neural networks — Keller Jordan et al., 2024 https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks 3. 4-bit Shampoo for Memory-Efficient Network Training — Sike Wang, Jia Li, Pan Zhou, Hua Huang, 2024 https://scholar.google.com/scholar?q=4-bit+Shampoo+for+Memory-Efficient+Network+Training 4. Combining axes preconditioners through Kronecker approximation for deep learning — Sai Surya Duvvuri, Fnu Devvrit, Rohan Anil, Cho-Jui Hsieh, Inderjit S. Dhillon, 2024 https://scholar.google.com/scholar?q=Combining+axes+preconditioners+through+Kronecker+approximation+for+deep+learning 5. No train no gain: Revisiting efficient training algorithms for transformer-based language models — Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, Matt J. Kusner, 2023 https://scholar.google.com/scholar?q=No+train+no+gain%3A+Revisiting+efficient+training+algorithms+for+transformer-based+language+models Interactive Visualization: SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam

Episode metadata supplied by the publisher feed · Published Aug 21, 2026

Embed this episode

NOW PLAYING

SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on August 21, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!