Abstract
Large language models (LLMs) are increasingly used for program verification, and yet little is known about how they reason about program semantics during this process. In this work, we focus on abstract interpretation based-reasoning for invariant generation and introduce two novel prompting strategies that aim to elicit such reasoning from LLMs. We evaluate these strategies across several state-of-the-art LLMs on 22 programs from the SV-COMP benchmark suite widely used in software verification. We analyze both the soundness of the generated invariants and the key thematic patterns in the models’ reasoning errors. This work aims to highlight new research opportunities at the intersection of LLMs and program verification for applying LLMs to verification tasks and advancing their reasoning capabilities in this application.
Author supplied keywords
Cite
CITATION STYLE
Mitchell, J., Zhou, C., Kim, B. H., & Wang, C. (2025). Understanding Formal Reasoning Failures in LLMs as Abstract Interpreters. In LMPL 2025 - Proceedings of the 1st ACM SIGPLAN International Workshop on Language Models and Programming Languages, Co-located with ICFP/SPLASH 2025 (pp. 71–83). Association for Computing Machinery, Inc. https://doi.org/10.1145/3759425.3763389
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.