DafnyBench and LLMs for Formal Verification episode artwork

EPISODE · Jun 17, 2026

DafnyBench and LLMs for Formal Verification

from AI Post Transformers

This episode explores DafnyBench, a benchmark for testing whether large language models can help with one of formal verification’s hardest practical bottlenecks: reconstructing the missing assertions and loop invariants that make Dafny programs verifiable. It explains how formal verification differs from ordinary testing and from theorem proving, and why the paper deliberately frames the task as restoring proof hints in existing verified programs rather than synthesizing correct software from scratch. The discussion digs into benchmark design, including the dataset of 782 single-file Dafny programs, the rule that models must infer both the content and placement of missing hints, and the importance of excluding shortcut tricks like disabling verification. It also highlights a crucial result nuance: 208 files already verify after hint removal, so the reported top score of about 67.8% is more informative when translated into genuine recovery performance on the subset that actually needs new annotations. Sources: 1. DafnyBench: A Benchmark for Formal Software Verification — Chloe Loughridge, Qinyi Sun, Seth Ahrenbach, Federico Cassano, Chuyue Sun, Ying Sheng, Anish Mudide, Md Rakib Hossain Misu, Nada Amin, Max Tegmark, 2024 http://arxiv.org/abs/2406.08467 2. Clover: Closed-Loop Verifiable Code Generation — Chuyue Sun, Ying Sheng, Oded Padon, Clark Barrett, 2024 https://scholar.google.com/scholar?q=Clover%3A+Closed-Loop+Verifiable+Code+Generation 3. Towards AI-Assisted Synthesis of Verified Dafny Methods — Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, 2024 https://scholar.google.com/scholar?q=Towards+AI-Assisted+Synthesis+of+Verified+Dafny+Methods 4. LeanDojo: Theorem Proving with Retrieval-Augmented Language Models — Kaiyu Yang et al., 2023 https://scholar.google.com/scholar?q=LeanDojo%3A+Theorem+Proving+with+Retrieval-Augmented+Language+Models 5. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code — Naman Jain et al., 2024 https://scholar.google.com/scholar?q=LiveCodeBench%3A+Holistic+and+Contamination+Free+Evaluation+of+Large+Language+Models+for+Code 6. Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification — Xu Xu et al., 2025 https://scholar.google.com/scholar?q=Local+Success+Does+Not+Compose%3A+Benchmarking+Large+Language+Models+for+Compositional+Formal+Verification 7. A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification — Norbert Tihanyi et al., 2023 https://scholar.google.com/scholar?q=A+New+Era+in+Software+Security%3A+Towards+Self-Healing+Software+via+Large+Language+Models+and+Formal+Verification 8. AI Post Transformers: LLM Agents Reason About Code Without Running It — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-15-llm-agents-reason-about-code-without-run-2a1876.mp3 9. AI Post Transformers: SkillsBench for Evaluating Agent Skills — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-14-skillsbench-for-evaluating-agent-skills-58bb1e.mp3

Episode metadata supplied by the publisher feed · Published Jun 17, 2026

Embed this episode

NOW PLAYING

DafnyBench and LLMs for Formal Verification

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 17, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!