EPISODE · May 7, 2026
Why LightGBM Made Boosted Trees Fast
from AI Post Transformers
This episode explores why LightGBM became a dominant tool for tabular machine learning by unpacking the algorithmic and systems ideas behind its speed. It explains how gradient boosting decision trees work, why split search becomes expensive on massive sparse datasets, and how LightGBM differs from neural-network-style training despite using gradient information. The discussion focuses on two core contributions: Gradient-based One-Side Sampling, which keeps high-gradient examples while subsampling easier ones without badly distorting split-gain estimates, and Exclusive Feature Bundling, which compresses sparse features by grouping columns that rarely activate together. Listeners would find it interesting for its clear account of how classical ideas like histograms, greedy tree growth, and graph coloring were combined into a highly practical system that reshaped real-world applications such as ranking, fraud detection, credit scoring, and forecasting. Sources: 1. Why LightGBM Made Boosted Trees Fast https://proceedings.neurips.cc/paper_files/paper/2017/file/6449f44a102fde848669bdd9eb6b76fa-Paper.pdf 2. Greedy Function Approximation: A Gradient Boosting Machine — Jerome H. Friedman, 2001 https://scholar.google.com/scholar?q=Greedy+Function+Approximation%3A+A+Gradient+Boosting+Machine 3. XGBoost: A Scalable Tree Boosting System — Tianqi Chen and Carlos Guestrin, 2016 https://scholar.google.com/scholar?q=XGBoost%3A+A+Scalable+Tree+Boosting+System 4. LightGBM: A Highly Efficient Gradient Boosting Decision Tree — Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, Tie-Yan Liu, 2017 https://scholar.google.com/scholar?q=LightGBM%3A+A+Highly+Efficient+Gradient+Boosting+Decision+Tree 5. CatBoost: Unbiased Boosting with Categorical Features — Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, Andrey Gulin, 2018 https://scholar.google.com/scholar?q=CatBoost%3A+Unbiased+Boosting+with+Categorical+Features 6. Feature Hashing for Large Scale Multitask Learning — Kilian Weinberger, Anirban Dasgupta, Josh Attenberg, John Langford, Alex Smola, 2009 https://scholar.google.com/scholar?q=Feature+Hashing+for+Large+Scale+Multitask+Learning 7. An Upper Bound for the Chromatic Number of a Graph and Its Application to Timetabling Problems — D. J. A. Welsh and M. B. Powell, 1967 https://scholar.google.com/scholar?q=An+Upper+Bound+for+the+Chromatic+Number+of+a+Graph+and+Its+Application+to+Timetabling+Problems 8. New Methods to Color the Vertices of a Graph — Daniel Brélaz, 1979 https://scholar.google.com/scholar?q=New+Methods+to+Color+the+Vertices+of+a+Graph 9. Worst Case Behavior of Graph Coloring Algorithms — David S. Johnson, 1974 https://scholar.google.com/scholar?q=Worst+Case+Behavior+of+Graph+Coloring+Algorithms 10. A Communication-Efficient Parallel Algorithm for Decision Tree — Qi Meng, Guolin Ke, Taifeng Wang, Wei Chen, Qiwei Ye, Zhi-Ming Ma, Tie-Yan Liu, 2016 https://scholar.google.com/scholar?q=A+Communication-Efficient+Parallel+Algorithm+for+Decision+Tree 11. Stochastic Gradient Boosting — Jerome H. Friedman, 2002 https://scholar.google.com/scholar?q=Stochastic+Gradient+Boosting 12. Parallel Boosted Regression Trees for Web Search Ranking — Stephen Tyree, Kilian Q. Weinberger, Kunal Agrawal, and Jennifer Paykin, 2011 https://scholar.google.com/scholar?q=Parallel+Boosted+Regression+Trees+for+Web+Search+Ranking 13. Best-First Decision Tree Learning — Haijian Shi, 2007 https://scholar.google.com/scholar?q=Best-First+Decision+Tree+Learning 14. GPU-Acceleration for Large-Scale Tree Boosting — Huan Zhang, Si Si, and Cho-Jui Hsieh, 2017 https://scholar.google.com/scholar?q=GPU-Acceleration+for+Large-Scale+Tree+Boosting 15. Implementing machine learning methods with complex survey data: Lessons learned on the impacts of accounting sampling weights in gradient boosting — authors not identified in the provided snippet, recent (2020s) https://scholar.google.com/scholar?q=Implementing+machine+learning+methods+with+complex+survey+data%3A+Lessons+learned+on+the+impacts+of+accounting+sampling+weights+in+gradient+boosting 16. Explainable boosting algorithms: sparse-group and interaction-aware variable selection in complex data — authors not identified in the provided snippet, recent (2020s) https://scholar.google.com/scholar?q=Explainable+boosting+algorithms%3A+sparse-group+and+interaction-aware+variable+selection+in+complex+data 17. Multi-objective optimization of performance and interpretability of tabular supervised machine learning models — authors not identified in the provided snippet, recent (2020s) https://scholar.google.com/scholar?q=Multi-objective+optimization+of+performance+and+interpretability+of+tabular+supervised+machine+learning+models 18. AI Post Transformers: Breiman's Two Cultures of Statistical Modeling — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-24-breimans-two-cultures-of-statistical-mod-71e49f.mp3 Interactive Visualization: Why LightGBM Made Boosted Trees Fast
Embed this episode
NOW PLAYING
Why LightGBM Made Boosted Trees Fast
No transcript for this episode yet
Similar Episodes
No similar episodes found.