Technical AGI Safety and Security Framework episode artwork

EPISODE · Jun 5, 2026

Technical AGI Safety and Security Framework

from AI Post Transformers

This episode explores DeepMind’s paper on technical AGI safety and security, focusing on how labs might prevent severe, humanity-scale harm before highly capable systems are deployed. It breaks down the paper’s core distinctions between misuse and misalignment, explains what the authors mean by Exceptional AGI and the no-human-ceiling assumption, and examines dangerous capability evaluations in areas like cyber, biology, persuasion, and self-proliferation. The discussion highlights the paper’s main argument that safety measures such as refusal training, jailbreak hardening, access controls, monitoring, anomaly detection, and model-weight security only matter if they are explicitly tied to capability thresholds that trigger real deployment restrictions. Listeners would find it interesting because it turns abstract AGI risk debates into a concrete governance and engineering framework for deciding when a model is too dangerous to release under normal conditions. Sources: 1. Technical AGI Safety and Security Framework https://arxiv.org/pdf/2504.01849 2. The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation — Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, et al., 2018 https://scholar.google.com/scholar?q=The+Malicious+Use+of+Artificial+Intelligence%3A+Forecasting%2C+Prevention%2C+and+Mitigation 3. Model evaluation for extreme risks — Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, et al., 2023 https://scholar.google.com/scholar?q=Model+evaluation+for+extreme+risks 4. Frontier AI Regulation: Managing Emerging Risks to Public Safety — Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, et al., 2023 https://scholar.google.com/scholar?q=Frontier+AI+Regulation%3A+Managing+Emerging+Risks+to+Public+Safety 5. Evaluating Frontier Models for Dangerous Capabilities — Mary Phuong, Matthew Aitchison, Elliot Catt, Victoria Krakovna, et al., 2024 https://scholar.google.com/scholar?q=Evaluating+Frontier+Models+for+Dangerous+Capabilities 6. Guidance on the Assurance of Machine Learning in Autonomous Systems (AMLAS) — Richard Hawkins, Colin Paterson, Chiara Picardi, Ibrahim Habli, et al., 2021 https://scholar.google.com/scholar?q=Guidance+on+the+Assurance+of+Machine+Learning+in+Autonomous+Systems+%28AMLAS%29 7. Safety Cases: How to Justify the Safety of Advanced AI Systems — Joshua Clymer, Nick Gabrieli, David Krueger, Thomas Larsen, 2024 https://scholar.google.com/scholar?q=Safety+Cases%3A+How+to+Justify+the+Safety+of+Advanced+AI+Systems 8. Safety case template for frontier AI: A cyber inability argument — Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett, Geoffrey Irving, et al., 2024 https://scholar.google.com/scholar?q=Safety+case+template+for+frontier+AI%3A+A+cyber+inability+argument 9. The BIG Argument for AI Safety Cases — Ibrahim Habli, Richard Hawkins, Colin Paterson, Mark Sujan, et al., 2025 https://scholar.google.com/scholar?q=The+BIG+Argument+for+AI+Safety+Cases 10. Safety cases: Justifying the safety of advanced AI systems — J. Clymer, N. Gabrieli, D. Krueger, and T. Larsen, 2024 https://scholar.google.com/scholar?q=Safety+cases%3A+Justifying+the+safety+of+advanced+AI+systems 11. AI control: Improving safety despite intentional subversion — R. Greenblatt, B. Shlegeris, K. Sachan, and F. Roger, 2024 https://scholar.google.com/scholar?q=AI+control%3A+Improving+safety+despite+intentional+subversion 12. Alignment faking in large language models — R. Greenblatt, C. Denison, B. Wright, F. Roger, M. MacDiarmid, S. Marks, J. Treutlein, T. Belonax, J. Chen, D. Duvenaud, et al., 2024 https://scholar.google.com/scholar?q=Alignment+faking+in+large+language+models 13. Towards evaluations-based safety cases for AI scheming — M. Balesni, M. Hobbhahn, D. Lindner, A. Meinke, T. Korbak, J. Clymer, B. Shlegeris, J. Scheurer, C. Stix, R. Shah, et al., 2024 https://scholar.google.com/scholar?q=Towards+evaluations-based+safety+cases+for+AI+scheming 14. Generative AI misuse: A taxonomy of tactics and insights from real-world data — N. Marchal, R. Xu, R. Elasmar, I. Gabriel, B. Goldberg, and W. Isaac, 2024 https://scholar.google.com/scholar?q=Generative+AI+misuse%3A+A+taxonomy+of+tactics+and+insights+from+real-world+data 15. Stress-Testing Capability Elicitation With Password-Locked Models — Ryan Greenblatt et al., 2024 https://scholar.google.com/scholar?q=Stress-Testing+Capability+Elicitation+With+Password-Locked+Models 16. Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models — Cameron Tice et al., 2024 https://scholar.google.com/scholar?q=Noise+Injection+Reveals+Hidden+Capabilities+of+Sandbagging+Language+Models 17. Benchmarking Misuse Mitigation Against Covert Adversaries — Davis Brown et al., 2025 https://scholar.google.com/scholar?q=Benchmarking+Misuse+Mitigation+Against+Covert+Adversaries 18. On scalable oversight with weak LLMs judging strong LLMs — Zachary Kenton et al., 2024 https://scholar.google.com/scholar?q=On+scalable+oversight+with+weak+LLMs+judging+strong+LLMs 19. Scaling Laws For Scalable Oversight — Joshua Engels et al., 2025 https://scholar.google.com/scholar?q=Scaling+Laws+For+Scalable+Oversight 20. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs — Kyle O'Brien et al., 2025 https://scholar.google.com/scholar?q=Deep+Ignorance%3A+Filtering+Pretraining+Data+Builds+Tamper-Resistant+Safeguards+into+Open-Weight+LLMs 21. AI Post Transformers: When LLM Judges Become Coin Flips — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-when-llm-judges-become-coin-flips-8b43ef.mp3 22. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3 23. AI Post Transformers: Split Personality Training Reveals Latent Knowledge — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-08-split-personality-training-reveals-laten-c84616.mp3 24. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3

Episode metadata supplied by the publisher feed · Published Jun 5, 2026

Embed this episode

NOW PLAYING

Technical AGI Safety and Security Framework

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 5, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!