EPISODE · Jul 27, 2026
Hilbert Operator Reveals What Networks Actually Learn
from AI Post Transformers
This episode examines HOPE (Hilbert Operator for Progressive Encoding), a structured pruning framework from Google DeepMind and UC Berkeley researchers that treats network compression as a diagnostic tool for understanding what deep networks actually learn, rather than just a deployment optimization. The discussion traces the approach's roots to the Information Bottleneck principle while carefully distinguishing HOPE's falsifiable measurement machinery from that unproven theory, and covers why magnitude-based pruning fails due to scale symmetry in batch-normalized networks, and how data-dependent pruning can quietly degrade long-tail class performance. The core innovation discussed is representing neurons as objects in a Hilbert space—comparing what function each neuron computes rather than the size of its weights—using only batch norm statistics already stored in a checkpoint, with no forward passes on real data and no hyperparameter tuning required. Listeners interested in interpretability, pruning theory, or the ongoing debate over why deep learning generalizes will find the hosts' back-and-forth on contested claims particularly engaging, as they push back on overstating the Information Bottleneck's explanatory power while crediting HOPE's mathematically rigorous, data-free approach to isolating a network's predictive core. Sources: 1. Hilbert Operator Reveals What Networks Actually Learn https://arxiv.org/pdf/2607.21366 2. Optimal Brain Damage — Yann LeCun, John S. Denker, Sara A. Solla, 1989 https://scholar.google.com/scholar?q=Optimal+Brain+Damage 3. The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks — Jonathan Frankle, Michael Carbin, 2019 https://scholar.google.com/scholar?q=The+Lottery+Ticket+Hypothesis%3A+Finding+Sparse%2C+Trainable+Neural+Networks 4. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015 https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network 5. SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot — Elias Frantar, Dan Alistarh, 2023 https://scholar.google.com/scholar?q=SparseGPT%3A+Massive+Language+Models+Can+Be+Accurately+Pruned+in+One-Shot 6. Priors for Infinite Networks (in Bayesian Learning for Neural Networks) — Radford M. Neal, 1996 https://scholar.google.com/scholar?q=Priors+for+Infinite+Networks+%28in+Bayesian+Learning+for+Neural+Networks%29 7. Neural Tangent Kernel: Convergence and Generalization in Neural Networks — Arthur Jacot, Franck Gabriel, Clément Hongler, 2018 https://scholar.google.com/scholar?q=Neural+Tangent+Kernel%3A+Convergence+and+Generalization+in+Neural+Networks 8. Fourier Neural Operator for Parametric Partial Differential Equations — Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, Anima Anandkumar, 2021 https://scholar.google.com/scholar?q=Fourier+Neural+Operator+for+Parametric+Partial+Differential+Equations 9. Measuring Statistical Dependence with Hilbert-Schmidt Norms — Arthur Gretton, Olivier Bousquet, Alex Smola, Bernhard Schölkopf, 2005 https://scholar.google.com/scholar?q=Measuring+Statistical+Dependence+with+Hilbert-Schmidt+Norms 10. What Do Compressed Deep Neural Networks Forget? — Sara Hooker, Aaron Courville, Gregory Clark, Yann Yannakakis, Kevin Murphy, 2019 https://scholar.google.com/scholar?q=What+Do+Compressed+Deep+Neural+Networks+Forget%3F 11. Towards Monosemanticity: Decomposing Language Models with Dictionary Learning — Trenton Bricken, Adly Templeton, Joshua Batson, et al. (Anthropic), 2023 https://scholar.google.com/scholar?q=Towards+Monosemanticity%3A+Decomposing+Language+Models+with+Dictionary+Learning 12. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, et al., 2022 https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models 13. Language Modeling Is Compression — Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, et al., 2023 https://scholar.google.com/scholar?q=Language+Modeling+Is+Compression 14. LLM-Pruner: On the Structural Pruning of Large Language Models — Xinyin Ma, Gongfan Fang, Xinchao Wang, 2023 https://scholar.google.com/scholar?q=LLM-Pruner%3A+On+the+Structural+Pruning+of+Large+Language+Models Interactive Visualization: Hilbert Operator Reveals What Networks Actually Learn
Embed this episode
NOW PLAYING
Hilbert Operator Reveals What Networks Actually Learn
No transcript for this episode yet
Similar Episodes
No similar episodes found.