EPISODE · Jul 17, 2026 · 1H 19M
Why AI Evaluations Are Broken and How to Fix Them (with David Manheim)
from Future of Life Institute Podcast · host Future of Life Institute
David Manheim is head of methodology at AI Evaluation Consensus. He joins the podcast to discuss how AI evaluations can become more reliable, transparent, and useful for decisions. We cover common failures such as unclear reporting, training to the test, benchmark saturation, and models changing behavior when they know they are being tested. The conversation also examines real-world tests, biosecurity, persuasion, forecasting, human oversight, and why even “normal” AI progress could be disruptive.LINKS:David Manheim websiteAI Evaluation Consensus StatementCHAPTERS:(00:00) Episode Preview(01:04) Evaluation consensus project(07:01) Evaluation awareness challenges(12:28) Reporting capabilities clearly(19:38) Benchmarks beyond humans(29:52) Proxies and biosecurity(42:01) Persuasion and democracy(53:59) Forecasting with AI(01:08:44) Oversight and disruption(01:16:42) Supporting better evalsPRODUCED BY:https://aipodcast.ingSOCIAL LINKS:Website: https://podcast.futureoflife.orgTwitter (FLI): https://x.com/FLI_orgTwitter (Gus): https://x.com/gusdockerLinkedIn: https://www.linkedin.com/company/future-of-life-institute/YouTube: https://www.youtube.com/channel/UC-rCCy3FQ-GItDimSR9lhzw/Apple: https://geo.itunes.apple.com/us/podcast/id1170991978Spotify: https://open.spotify.com/show/2Op1WO3gwVwCrYHg4eoGyP
Embed this episode
NOW PLAYING
Why AI Evaluations Are Broken and How to Fix Them (with David Manheim)
No transcript for this episode yet
Similar Episodes
Similar Podcasts
No similar podcasts found.