Can LLMs Enable Verification in Mainstream Programming?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shefer, Aleksandr, Engel, Igor, Alekseev, Stanislav, Berezun, Daniil, Verbitskaia, Ekaterina, Podkopaev, Anton
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908273566810112
author Shefer, Aleksandr
Engel, Igor
Alekseev, Stanislav
Berezun, Daniil
Verbitskaia, Ekaterina
Podkopaev, Anton
author_facet Shefer, Aleksandr
Engel, Igor
Alekseev, Stanislav
Berezun, Daniil
Verbitskaia, Ekaterina
Podkopaev, Anton
contents Although formal methods are capable of producing reliable software, they have seen minimal adoption in everyday programming. Automatic code generation using large language models is becoming increasingly widespread, but it rarely considers producing strong correctness guarantees. In this study, we explore the ability of LLMs to produce verified code in three verification languages (Dafny, Nagini, and Verus). To do so, we use manually curated datasets derived from the state-ofthe-art Python benchmark, HumanEval. We also assess what types of information are sufficient to achieve good-quality results.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14183
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can LLMs Enable Verification in Mainstream Programming?
Shefer, Aleksandr
Engel, Igor
Alekseev, Stanislav
Berezun, Daniil
Verbitskaia, Ekaterina
Podkopaev, Anton
Software Engineering
Artificial Intelligence
Programming Languages
Although formal methods are capable of producing reliable software, they have seen minimal adoption in everyday programming. Automatic code generation using large language models is becoming increasingly widespread, but it rarely considers producing strong correctness guarantees. In this study, we explore the ability of LLMs to produce verified code in three verification languages (Dafny, Nagini, and Verus). To do so, we use manually curated datasets derived from the state-ofthe-art Python benchmark, HumanEval. We also assess what types of information are sufficient to achieve good-quality results.
title Can LLMs Enable Verification in Mainstream Programming?
topic Software Engineering
Artificial Intelligence
Programming Languages
url https://arxiv.org/abs/2503.14183