EPISODE · Jun 15, 2026
Dafny for Trustworthy AI Code Generation
from AI Post Transformers
This episode explores a 2025 paper on using Dafny as a hidden, verification-aware intermediate language for AI code generation, where a model first produces a formal specification and verified implementation before compiling it into ordinary Python. It examines the paper’s central trust claim: formal verification can prove that the generated code satisfies the hidden spec, but it cannot prove that the hidden spec actually matches the user’s intent, making spec alignment a separate and critical failure point. The discussion uses the paper’s fibfib example and HumanEval results to unpack that distinction, noting that the Dafny-only pipeline trails direct Python generation, while the best reported score comes only after falling back to unverified Python when the verification loop fails to converge. Listeners would find it interesting because it gives a concrete, nuanced look at where AI coding assistants can become more reliable, where the guarantees stop, and why neuro-symbolic workflows may matter most for tightly specified code like algorithms, parsers, and protocol logic. Sources: 1. Dafny as Verification-Aware Intermediate Language for Code Generation — Yue Chen Li, Stefan Zetzsche, Siva Somayyajula, 2025 http://arxiv.org/abs/2501.06283 2. Dafny: An Automatic Program Verifier for Functional Correctness — K. Rustan M. Leino, 2010 https://scholar.google.com/scholar?q=Dafny%3A+An+Automatic+Program+Verifier+for+Functional+Correctness 3. Towards AI-Assisted Synthesis of Verified Dafny Methods — Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, 2024 https://scholar.google.com/scholar?q=Towards+AI-Assisted+Synthesis+of+Verified+Dafny+Methods 4. Clover: Closed-Loop Verifiable Code Generation — Chuyue Sun, Ying Sheng, Oded Padon, Clark Barrett, 2024 https://scholar.google.com/scholar?q=Clover%3A+Closed-Loop+Verifiable+Code+Generation 5. VerMCTS: Synthesizing Multi-Step Programs using a Verifier, a Large Language Model, and Tree Search — David Brandfonbrener, Simon Henniger, Sibi Raja, Tarun Prasad, Chloe Loughridge, Federico Cassano, Sabrina Ruixin Hu, Jianang Yang, William E. Byrd, Robert Zinkov, Nada Amin, 2024 https://scholar.google.com/scholar?q=VerMCTS%3A+Synthesizing+Multi-Step+Programs+using+a+Verifier%2C+a+Large+Language+Model%2C+and+Tree+Search 6. Laurel: Generating Dafny Assertions Using Large Language Models — Eric Mugnier, Emmanuel Anaya Gonzalez, Ranjit Jhala, Nadia Polikarpova, Yuanyuan Zhou, 2024 https://scholar.google.com/scholar?q=Laurel%3A+Generating+Dafny+Assertions+Using+Large+Language+Models 7. DafnyBench: A Benchmark for Formal Software Verification — Chloe Loughridge et al., 2024 https://scholar.google.com/scholar?q=DafnyBench%3A+A+Benchmark+for+Formal+Software+Verification 8. Evaluating Large Language Models Trained on Code — Mark Chen et al., 2021 https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code 9. Baking for Dafny: A CakeML Backend for Dafny — Daniel Nezamabadi, Magnus Myreen, 2025 https://scholar.google.com/scholar?q=Baking+for+Dafny%3A+A+CakeML+Backend+for+Dafny 10. Intent-aligned Formal Specification Synthesis via Traceable Refinement — Zhe Ye et al., 2026 https://scholar.google.com/scholar?q=Intent-aligned+Formal+Specification+Synthesis+via+Traceable+Refinement 11. Combining LLM Code Generation with Formal Specifications and Reactive Program Synthesis — William Murphy et al., 2024 https://scholar.google.com/scholar?q=Combining+LLM+Code+Generation+with+Formal+Specifications+and+Reactive+Program+Synthesis 12. StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback — Shihan Dou et al., 2024 https://scholar.google.com/scholar?q=StepCoder%3A+Improve+Code+Generation+with+Reinforcement+Learning+from+Compiler+Feedback 13. InspectCoder: Dynamic Analysis-Enabled Self Repair through interactive LLM-Debugger Collaboration — Yunkun Wang et al., 2025 https://scholar.google.com/scholar?q=InspectCoder%3A+Dynamic+Analysis-Enabled+Self+Repair+through+interactive+LLM-Debugger+Collaboration 14. FormalSpecCpp: A Dataset of C++ Formal Specifications created using LLMs — Madhurima Chakraborty et al., 2025 https://scholar.google.com/scholar?q=FormalSpecCpp%3A+A+Dataset+of+C%2B%2B+Formal+Specifications+created+using+LLMs 15. AI Post Transformers: From Natural Language to Verified Dafny Code — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-14-from-natural-language-to-verified-dafny-8abed9.mp3 16. AI Post Transformers: Program Synthesis with Large Language Models — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-20-program-synthesis-with-large-language-mo-b962ec.mp3 17. AI Post Transformers: Generative File Systems: Replacing Code with Formal Specifications — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-18-generative-file-systems-replacing-code-w-414029.mp3
Embed this episode
NOW PLAYING
Dafny for Trustworthy AI Code Generation
No transcript for this episode yet
Similar Episodes
No similar episodes found.