EPISODE · Jun 14, 2026
From Natural Language to Verified Dafny Code
from AI Post Transformers
This episode explores a 2026 study on turning long natural-language programming problems into Dafny code that can be formally verified, asking whether AI systems can produce code that is not just fluent but provably correct. It explains how Dafny uses preconditions, postconditions, loop invariants, and proof obligations, and why weak specifications can lead to vacuous “verified” programs that still fail to capture the real task. The discussion highlights the paper’s NL2VC-60 benchmark of hand-written verified solutions to UVa-style algorithm problems, along with experiments comparing plain prompting, signature-guided prompting, and self-healing loops that revise code using verifier feedback and additional uDebug testing. Listeners would find it interesting because it gets at the core trust problem in AI coding: whether formal methods can make generated software more reliable, and where the real bottleneck remains the human effort required to write strong specifications. Sources: 1. From Natural Language to Verified Code: Toward AI Assisted Problem-to-Code Generation with Dafny-Based Formal Verification — Md Erfan, Md Kamal Hossain Chowdhury, Ahmed Ryan, Md Rayhanur Rahman, 2026 http://arxiv.org/abs/2604.22601 2. Dafny: An Automatic Program Verifier for Functional Correctness — K. Rustan M. Leino, 2010 https://scholar.google.com/scholar?q=Dafny%3A+An+Automatic+Program+Verifier+for+Functional+Correctness 3. seL4: Formal Verification of an Operating-System Kernel — Gerwin Klein, Kevin Elphinstone, Gernot Heiser, June Andronick, David Cock, et al., 2009 https://scholar.google.com/scholar?q=seL4%3A+Formal+Verification+of+an+Operating-System+Kernel 4. Formal verification of a realistic compiler — Xavier Leroy, 2009 https://scholar.google.com/scholar?q=Formal+verification+of+a+realistic+compiler 5. Modularity, Code Specialization, and Zero-Cost Abstractions for Program Verification — Son Ho, Aymeric Fromherz, Jonathan Protzenko, 2021 https://scholar.google.com/scholar?q=Modularity%2C+Code+Specialization%2C+and+Zero-Cost+Abstractions+for+Program+Verification 6. Towards AI-Assisted Synthesis of Verified Dafny Methods — Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, 2024 https://scholar.google.com/scholar?q=Towards+AI-Assisted+Synthesis+of+Verified+Dafny+Methods 7. DafnyBench: A Benchmark for Formal Software Verification — Chloe Loughridge et al., 2024 https://scholar.google.com/scholar?q=DafnyBench%3A+A+Benchmark+for+Formal+Software+Verification 8. Can LLMs Enable Verification in Mainstream Programming? — Aleksandr Shefer, Igor Engel, Stanislav Alekseev, Daniil Berezun, Ekaterina Verbitskaia, Anton Podkopaev, 2025 https://scholar.google.com/scholar?q=Can+LLMs+Enable+Verification+in+Mainstream+Programming%3F 9. Dafny as Verification-Aware Intermediate Language for Code Generation — Yue Chen Li, Stefan Zetzsche, Siva Somayyajula, 2025 https://scholar.google.com/scholar?q=Dafny+as+Verification-Aware+Intermediate+Language+for+Code+Generation 10. ATLAS: Automated Toolkit for Large-Scale Verified Code Synthesis — Mantas Baksys et al., 2025 https://scholar.google.com/scholar?q=ATLAS%3A+Automated+Toolkit+for+Large-Scale+Verified+Code+Synthesis 11. DafnyPro: LLM-Assisted Automated Verification for Dafny Programs — Debangshu Banerjee, Olivier Bouissou, Stefan Zetzsche, 2026 https://scholar.google.com/scholar?q=DafnyPro%3A+LLM-Assisted+Automated+Verification+for+Dafny+Programs 12. Neuro Symbolic Reasoning for Planning: Counterexample Guided Inductive Synthesis using Large Language Models and Satisfiability Solving — Sumit Kumar Jha et al., 2023 https://scholar.google.com/scholar?q=Neuro+Symbolic+Reasoning+for+Planning%3A+Counterexample+Guided+Inductive+Synthesis+using+Large+Language+Models+and+Satisfiability+Solving 13. Property-Guided LLM Program Synthesis for Planning — Andre G. Pereira, Augusto B. Correa, Jendrik Seipp, 2026 https://scholar.google.com/scholar?q=Property-Guided+LLM+Program+Synthesis+for+Planning 14. Finding Inductive Loop Invariants using Large Language Models — Adharsh Kamath et al., 2023 https://scholar.google.com/scholar?q=Finding+Inductive+Loop+Invariants+using+Large+Language+Models 15. LLM For Loop Invariant Generation and Fixing: How Far Are We? — Mostafijur Rahman Akhond, Saikat Chakraborty, Gias Uddin, 2025 https://scholar.google.com/scholar?q=LLM+For+Loop+Invariant+Generation+and+Fixing%3A+How+Far+Are+We%3F 16. Type-Constrained Code Generation with Language Models — Niels Mundler et al., 2025 https://scholar.google.com/scholar?q=Type-Constrained+Code+Generation+with+Language+Models 17. Invariant-based Program Repair — Omar I. Al-Bataineh, 2024 https://scholar.google.com/scholar?q=Invariant-based+Program+Repair 18. AI Post Transformers: Program Synthesis with Large Language Models — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-20-program-synthesis-with-large-language-mo-b962ec.mp3 19. AI Post Transformers: SGLang for Faster Structured LLM Programs — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-06-sglang-for-faster-structured-llm-program-c59f1c.mp3 20. AI Post Transformers: SkillsBench for Evaluating Agent Skills — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-14-skillsbench-for-evaluating-agent-skills-58bb1e.mp3
Embed this episode
NOW PLAYING
From Natural Language to Verified Dafny Code
No transcript for this episode yet
Similar Episodes
No similar episodes found.