AdamW: Decoupled Weight Decay Regularization for Adaptive Gradient Algorithms episode artwork

EPISODE · Aug 8, 2025

AdamW: Decoupled Weight Decay Regularization for Adaptive Gradient Algorithms

from AI Post Transformers

The paper "Decoupled Weight Decay Regularization" by Ilya Loshchilov and Frank Hutter (2017) introduced what we know in Pytorch now as AdamW. This academic paper explores the differences between L2 regularization and weight decay regularization in optimizing neural networks, particularly focusing on adaptive gradient algorithms like Adam. The authors demonstrate that while these two regularization methods are equivalent for standard stochastic gradient descent (SGD), they are not equivalent for Adam, leading to suboptimal generalization performance in common Adam implementations. They propose a decoupled weight decay method, termed AdamW, which substantially improves Adam's generalization and allows its performance to rival SGD with momentum on image classification tasks. The paper also introduces normalized weight decay and discusses the integration of warm restarts for improved performance and hyperparameter tuning.

Episode metadata supplied by the publisher feed · Published Aug 8, 2025

Embed this episode

NOW PLAYING

AdamW: Decoupled Weight Decay Regularization for Adaptive Gradient Algorithms

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on August 8, 2025.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!