Can LLMs Enable Verification in Mainstream Programming?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908273566810112 |
|---|---|
| author | Shefer, Aleksandr Engel, Igor Alekseev, Stanislav Berezun, Daniil Verbitskaia, Ekaterina Podkopaev, Anton |
| author_facet | Shefer, Aleksandr Engel, Igor Alekseev, Stanislav Berezun, Daniil Verbitskaia, Ekaterina Podkopaev, Anton |
| contents | Although formal methods are capable of producing reliable software, they have seen minimal adoption in everyday programming. Automatic code generation using large language models is becoming increasingly widespread, but it rarely considers producing strong correctness guarantees. In this study, we explore the ability of LLMs to produce verified code in three verification languages (Dafny, Nagini, and Verus). To do so, we use manually curated datasets derived from the state-ofthe-art Python benchmark, HumanEval. We also assess what types of information are sufficient to achieve good-quality results. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_14183 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Can LLMs Enable Verification in Mainstream Programming? Shefer, Aleksandr Engel, Igor Alekseev, Stanislav Berezun, Daniil Verbitskaia, Ekaterina Podkopaev, Anton Software Engineering Artificial Intelligence Programming Languages Although formal methods are capable of producing reliable software, they have seen minimal adoption in everyday programming. Automatic code generation using large language models is becoming increasingly widespread, but it rarely considers producing strong correctness guarantees. In this study, we explore the ability of LLMs to produce verified code in three verification languages (Dafny, Nagini, and Verus). To do so, we use manually curated datasets derived from the state-ofthe-art Python benchmark, HumanEval. We also assess what types of information are sufficient to achieve good-quality results. |
| title | Can LLMs Enable Verification in Mainstream Programming? |
| topic | Software Engineering Artificial Intelligence Programming Languages |
| url | https://arxiv.org/abs/2503.14183 |