Saved in:
Bibliographic Details
Main Authors: Cai, Yufan, Hou, Zhe, Luan, Xiaokun, Baena, David Miguel Sanan, Lin, Yun, Sun, Jun, Dong, Jin Song
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2406.18616
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909232288235520
author Cai, Yufan
Hou, Zhe
Luan, Xiaokun
Baena, David Miguel Sanan
Lin, Yun
Sun, Jun
Dong, Jin Song
author_facet Cai, Yufan
Hou, Zhe
Luan, Xiaokun
Baena, David Miguel Sanan
Lin, Yun
Sun, Jun
Dong, Jin Song
contents Program refinement involves correctness-preserving transformations from formal high-level specification statements into executable programs. Traditional verification tool support for program refinement is highly interactive and lacks automation. On the other hand, the emergence of large language models (LLMs) enables automatic code generations from informal natural language specifications. However, code generated by LLMs is often unreliable. Moreover, the opaque procedure from specification to code provided by LLM is an uncontrolled black box. We propose LLM4PR, a tool that combines formal program refinement techniques with informal LLM-based methods to (1) transform the specification to preconditions and postconditions, (2) automatically build prompts based on refinement calculus, (3) interact with LLM to generate code, and finally, (4) verify that the generated code satisfies the conditions of refinement calculus, thus guaranteeing the correctness of the code. We have implemented our tool using GPT4, Coq, and Coqhammer, and evaluated it on the HumanEval and EvalPlus datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18616
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Large Language Model Aided Program Refinement
Cai, Yufan
Hou, Zhe
Luan, Xiaokun
Baena, David Miguel Sanan
Lin, Yun
Sun, Jun
Dong, Jin Song
Software Engineering
Artificial Intelligence
Computation and Language
K.6.3
Program refinement involves correctness-preserving transformations from formal high-level specification statements into executable programs. Traditional verification tool support for program refinement is highly interactive and lacks automation. On the other hand, the emergence of large language models (LLMs) enables automatic code generations from informal natural language specifications. However, code generated by LLMs is often unreliable. Moreover, the opaque procedure from specification to code provided by LLM is an uncontrolled black box. We propose LLM4PR, a tool that combines formal program refinement techniques with informal LLM-based methods to (1) transform the specification to preconditions and postconditions, (2) automatically build prompts based on refinement calculus, (3) interact with LLM to generate code, and finally, (4) verify that the generated code satisfies the conditions of refinement calculus, thus guaranteeing the correctness of the code. We have implemented our tool using GPT4, Coq, and Coqhammer, and evaluated it on the HumanEval and EvalPlus datasets.
title Towards Large Language Model Aided Program Refinement
topic Software Engineering
Artificial Intelligence
Computation and Language
K.6.3
url https://arxiv.org/abs/2406.18616