EPISODE · Oct 9, 2025 · 12 MIN
Learning dynamics of LLM finetuning
from Best AI papers explained · host Enoch H. Kang
This academic paper presents a novel framework for understanding the evolution of Large Language Models (LLMs) during finetuning by analyzing their learning dynamics from a dynamical perspective, contrasting with previous approaches focused on training targets or end-states. The authors formalize the change in model prediction using a decomposition into three key terms, which adapts to various finetuning algorithms like Supervised Finetuning (SFT) and Direct Preference Optimization (DPO). A significant finding is the "squeezing effect" caused by negative gradients during preference tuning, which reduces the confidence of most responses and is especially pronounced when the model is already confident or finetuning is off-policy. The framework is validated through experiments on both the MNIST dataset and LLM finetuning, demonstrating its ability to explain counter-intuitive phenomena like the confidence decay observed in DPO. Finally, the research inspires a simple yet effective method to improve alignment performance by mitigating the harmful aspects of the squeezing effect.
What this episode covers
This academic paper presents a novel framework for understanding the evolution of Large Language Models (LLMs) during finetuning by analyzing their learning dynamics from a dynamical perspective, contrasting with previous approaches focused on training targets or end-states. The authors formalize the change in model prediction using a decomposition into three key terms, which adapts to various finetuning algorithms like Supervised Finetuning (SFT) and Direct Preference Optimization (DPO). A significant finding is the "squeezing effect" caused by negative gradients during preference tuning, which reduces the confidence of most responses and is especially pronounced when the model is already confident or finetuning is off-policy. The framework is validated through experiments on both the MNIST dataset and LLM finetuning, demonstrating its ability to explain counter-intuitive phenomena like the confidence decay observed in DPO. Finally, the research inspires a simple yet effective method to improve alignment performance by mitigating the harmful aspects of the squeezing effect.
NOW PLAYING
Learning dynamics of LLM finetuning
No transcript for this episode yet
Similar Episodes
Mar 31, 2026 ·54m
Mar 27, 2026 ·14m
Mar 24, 2026 ·42m
Mar 20, 2026 ·42m
Mar 17, 2026 ·41m
Mar 13, 2026 ·44m