Can ChatGPT support software verification?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Janßen, Christian, Richter, Cedric, Wehrheim, Heike
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911209630990336
author Janßen, Christian
Richter, Cedric
Wehrheim, Heike
author_facet Janßen, Christian
Richter, Cedric
Wehrheim, Heike
contents Large language models have become increasingly effective in software engineering tasks such as code generation, debugging and repair. Language models like ChatGPT can not only generate code, but also explain its inner workings and in particular its correctness. This raises the question whether we can utilize ChatGPT to support formal software verification. In this paper, we take some first steps towards answering this question. More specifically, we investigate whether ChatGPT can generate loop invariants. Loop invariant generation is a core task in software verification, and the generation of valid and useful invariants would likely help formal verifiers. To provide some first evidence on this hypothesis, we ask ChatGPT to annotate 106 C programs with loop invariants. We check validity and usefulness of the generated invariants by passing them to two verifiers, Frama-C and CPAchecker. Our evaluation shows that ChatGPT is able to produce valid and useful invariants allowing Frama-C to verify tasks that it could not solve before. Based on our initial insights, we propose ways of combining ChatGPT (or large language models in general) and software verifiers, and discuss current limitations and open issues.
format Preprint
id arxiv_https___arxiv_org_abs_2311_02433
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Can ChatGPT support software verification?
Janßen, Christian
Richter, Cedric
Wehrheim, Heike
Software Engineering
Artificial Intelligence
Formal Languages and Automata Theory
Machine Learning
Logic in Computer Science
Large language models have become increasingly effective in software engineering tasks such as code generation, debugging and repair. Language models like ChatGPT can not only generate code, but also explain its inner workings and in particular its correctness. This raises the question whether we can utilize ChatGPT to support formal software verification. In this paper, we take some first steps towards answering this question. More specifically, we investigate whether ChatGPT can generate loop invariants. Loop invariant generation is a core task in software verification, and the generation of valid and useful invariants would likely help formal verifiers. To provide some first evidence on this hypothesis, we ask ChatGPT to annotate 106 C programs with loop invariants. We check validity and usefulness of the generated invariants by passing them to two verifiers, Frama-C and CPAchecker. Our evaluation shows that ChatGPT is able to produce valid and useful invariants allowing Frama-C to verify tasks that it could not solve before. Based on our initial insights, we propose ways of combining ChatGPT (or large language models in general) and software verifiers, and discuss current limitations and open issues.
title Can ChatGPT support software verification?
topic Software Engineering
Artificial Intelligence
Formal Languages and Automata Theory
Machine Learning
Logic in Computer Science
url https://arxiv.org/abs/2311.02433