When LLM Judges Become Coin Flips episode artwork

EPISODE · May 9, 2026

When LLM Judges Become Coin Flips

from AI Post Transformers

This episode explores a March 2026 paper arguing that LLM-based judges are an unreliable way to measure jailbreak success and adversarial robustness. It explains how modern safety evaluations rely on judge models to score harmful outputs, then walks through why those judges can break under attack shift, model shift, and data shift, sometimes degrading to near coin-flip reliability. The discussion connects this critique to benchmarks such as MT-Bench, HarmBench, and StrongREJECT, and examines how weaknesses in the judging pipeline can inflate or distort reported attack success rates. Listeners would find it interesting because it challenges whether many headline jailbreak results are exposing real model failures or simply failures in the grading system. Sources: 1. A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness — Leo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami, Gauthier Gidel, Stephan Günnemann, 2026 http://arxiv.org/abs/2603.06594 2. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zi Lin, Zhuohan Li, Joseph E. Gonzalez, Ion Stoica and others, 2023 https://scholar.google.com/scholar?q=Judging+LLM-as-a-Judge+with+MT-Bench+and+Chatbot+Arena 3. Large Language Models are not Fair Evaluators — Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, Zhifang Sui, 2023 https://scholar.google.com/scholar?q=Large+Language+Models+are+not+Fair+Evaluators 4. An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Models are Task-specific Classifiers — Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, Tiejun Zhao, 2024 https://scholar.google.com/scholar?q=An+Empirical+Study+of+LLM-as-a-Judge+for+LLM+Evaluation%3A+Fine-tuned+Judge+Models+are+Task-specific+Classifiers 5. JudgeBench: A Benchmark for Evaluating LLM-based Judges — Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y. Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, Ion Stoica, 2024 https://scholar.google.com/scholar?q=JudgeBench%3A+A+Benchmark+for+Evaluating+LLM-based+Judges 6. Universal and Transferable Adversarial Attacks on Aligned Language Models — Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, Matt Fredrikson, 2023 https://scholar.google.com/scholar?q=Universal+and+Transferable+Adversarial+Attacks+on+Aligned+Language+Models 7. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal — Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, Dan Hendrycks, 2024 https://scholar.google.com/scholar?q=HarmBench%3A+A+Standardized+Evaluation+Framework+for+Automated+Red+Teaming+and+Robust+Refusal 8. A StrongREJECT for Empty Jailbreaks — Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, Sam Toyer, 2024 https://scholar.google.com/scholar?q=A+StrongREJECT+for+Empty+Jailbreaks 9. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models — Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tramèr, Hamed Hassani, Eric Wong, 2024 https://scholar.google.com/scholar?q=JailbreakBench%3A+An+Open+Robustness+Benchmark+for+Jailbreaking+Large+Language+Models 10. Dataset Shift in Machine Learning — Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, Neil D. Lawrence (editors), 2008 https://scholar.google.com/scholar?q=Dataset+Shift+in+Machine+Learning 11. Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift — Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D. Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, Jasper Snoek, 2019 https://scholar.google.com/scholar?q=Can+You+Trust+Your+Model%27s+Uncertainty%3F+Evaluating+Predictive+Uncertainty+Under+Dataset+Shift 12. Measuring Robustness to Natural Distribution Shifts in Image Classification — Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, Ludwig Schmidt, 2020 https://scholar.google.com/scholar?q=Measuring+Robustness+to+Natural+Distribution+Shifts+in+Image+Classification 13. WILDS: A Benchmark of in-the-Wild Distribution Shifts — Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Percy Liang and others, 2021 https://scholar.google.com/scholar?q=WILDS%3A+A+Benchmark+of+in-the-Wild+Distribution+Shifts 14. LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge — Shuaizhi Li, Chenxu Xu, Jiazhu Wang, Xianyu Gong, Cheng Chen, Jun Zhang, Junjie Wang, Kit Lam, and Shouling Ji, 2025 https://scholar.google.com/scholar?q=LLMs+Cannot+Reliably+Judge+%28Yet%3F%29%3A+A+Comprehensive+Assessment+on+the+Robustness+of+LLM-as-a-Judge 15. Confusion Is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs — Yixin Yan, Shichao Sun, Zhen Wang, Yifan Lin, Zeyu Duan, Zhenzhen Zheng, Mingyu Liu, Zhenfei Yin, and Jie Zhang, 2025 https://scholar.google.com/scholar?q=Confusion+Is+the+Final+Barrier%3A+Rethinking+Jailbreak+Evaluation+and+Investigating+the+Real+Misuse+Threat+of+LLMs 16. Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges — Francisco Eiras et al., 2025 https://scholar.google.com/scholar?q=Know+Thy+Judge%3A+On+the+Robustness+Meta-Evaluation+of+LLM+Safety+Judges 17. Comparison Requires Valid Measurement: Rethinking Attack Success Rate Comparisons in AI Red Teaming — Alexandra Chouldechova, A. Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, Hanna Wallach, 2025/2026 https://scholar.google.com/scholar?q=Comparison+Requires+Valid+Measurement%3A+Rethinking+Attack+Success+Rate+Comparisons+in+AI+Red+Teaming 18. How to Correctly Report LLM-as-a-Judge Evaluations — Chungpa Lee et al., 2025 https://scholar.google.com/scholar?q=How+to+Correctly+Report+LLM-as-a-Judge+Evaluations 19. Embracing Ambiguity: Bayesian Nonparametrics and Stakeholder Participation for Ambiguity-Aware Safety Evaluation — Yanan Long, 2025/2026 https://scholar.google.com/scholar?q=Embracing+Ambiguity%3A+Bayesian+Nonparametrics+and+Stakeholder+Participation+for+Ambiguity-Aware+Safety+Evaluation 20. AI Post Transformers: Multidimensional Safety Evaluation of Frontier AI Models — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/multidimensional-safety-evaluation-of-frontier-ai-models/ 21. AI Post Transformers: LLM Benchmark Robustness to Linguistic Variation — Hal Turing & Dr. Ada Shannon, 2025 https://podcast.do-not-panic.com/episodes/llm-benchmark-robustness-to-linguistic-variation/ 22. AI Post Transformers: AI Agent Traps and Prompt Injection — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-02-ai-agent-traps-and-prompt-injection-7ce4ba.mp3 Interactive Visualization: When LLM Judges Become Coin Flips

Episode metadata supplied by the publisher feed · Published May 9, 2026

Embed this episode

NOW PLAYING

When LLM Judges Become Coin Flips

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 9, 2026.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!