EPISODE · Jul 6, 2026
Discretizing Reward Models for RL Alignment
from AI Post Transformers
This episode explores the paper Discretizing Reward Models and its argument that smooth decimal reward scores can be misleading for reinforcement learning alignment, because policies learn to exploit tiny, often meaningless differences instead of genuine quality. It explains why reward models are used for fuzzy goals like helpfulness and honesty, then digs into reward hacking, equivalence classes of equally valid answers, and the distinction between a model’s ability to separate good from bad responses versus its tendency to invent rankings among ties. The discussion also covers benchmarks such as the Ties setting and the paper’s core proposal: replacing continuous scores with a small number of ordinal reward buckets built from uncertainty estimates, pairwise equivalence judgments, and hierarchical clustering. Listeners would find it interesting because it connects an abstract modeling choice to a practical alignment problem facing modern language-model training, while also examining why the field currently seems more convinced by the diagnosis than by large-scale adoption of this exact fix. Sources: 1. Discretizing Reward Models — Vijay Viswanathan, Shiqi Wang, Devamanyu Hazarika, Chirag Nagpal, Tongshuang Wu, Graham Neubig, Yuning Mao, 2026 http://arxiv.org/abs/2606.21795 2. Deep Reinforcement Learning from Human Preferences — Paul Christiano, Jan Leike, Tom B. Brown, Shane Legg, Dario Amodei, 2017 https://scholar.google.com/scholar?q=Deep+Reinforcement+Learning+from+Human+Preferences 3. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback — Stephen Casper, Xander Davies, Claudia Shi, Jeremy Scheurer, Dylan Hadfield-Menell, et al., 2023 https://scholar.google.com/scholar?q=Open+Problems+and+Fundamental+Limitations+of+Reinforcement+Learning+from+Human+Feedback 4. RewardBench 2: Advancing Reward Model Evaluation — Saumya Malik, Valentina Pyatkin, Sander Land, Nathan Lambert, Noah A. Smith, Hannaneh Hajishirzi, 2025 https://scholar.google.com/scholar?q=RewardBench+2%3A+Advancing+Reward+Model+Evaluation 5. Discretizing Reward Models — Vijay Viswanathan, Shiqi Wang, Devamanyu Hazarika, Chirag Nagpal, Tongshuang Wu, Graham Neubig, Yuning Mao, 2026 https://scholar.google.com/scholar?q=Discretizing+Reward+Models 6. How to Evaluate Reward Models for RLHF — Evan Frick, Tianle Li, Connor Chen, Wei-Lin Chiang, Anastasios Nikolas Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph Gonzalez, Ion Stoica, 2024 https://scholar.google.com/scholar?q=How+to+Evaluate+Reward+Models+for+RLHF 7. What Makes a Reward Model a Good Teacher? An Optimization Perspective — Noam Razin, Zixuan Wang, Hubert Strauss, Stanley Wei, Jason D. Lee, Sanjeev Arora, 2025 https://scholar.google.com/scholar?q=What+Makes+a+Reward+Model+a+Good+Teacher%3F+An+Optimization+Perspective 8. The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models — Yanjun Chen, Dawei Zhu, Yirong Sun, Xinghao Chen, Wei Zhang, Xiaoyu Shen, 2024 https://scholar.google.com/scholar?q=The+Accuracy+Paradox+in+RLHF%3A+When+Better+Reward+Models+Don%27t+Yield+Better+Language+Models 9. Validating LLM-as-a-Judge Systems under Rating Indeterminacy — Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna Wallach, Zhiwei Steven Wu, Alexandra Chouldechova, 2025 https://scholar.google.com/scholar?q=Validating+LLM-as-a-Judge+Systems+under+Rating+Indeterminacy 10. Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback — Amirhossein Afsharrad, Ruida Zhou, Luca Viano, Sanjay Lall, Mohammad Ghavamzadeh, 2026 https://scholar.google.com/scholar?q=Beyond+Binary+Preferences%3A+A+Principled+Framework+for+Reward+Modeling+with+Ordinal+Feedback 11. Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts — Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, Tong Zhang, 2024 https://scholar.google.com/scholar?q=Interpretable+Preferences+via+Multi-Objective+Reward+Modeling+and+Mixture-of-Experts 12. AI Post Transformers: Split Personality Training Reveals Latent Knowledge — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-08-split-personality-training-reveals-laten-c84616.mp3 13. AI Post Transformers: Robots Need More Than VLAs and World Models — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-10-robots-need-more-than-vlas-and-world-mod-cdab8b.mp3
Embed this episode
NOW PLAYING
Discretizing Reward Models for RL Alignment
No transcript for this episode yet
Similar Episodes
No similar episodes found.