Automated Proof Generation for Rust Code via Self-Evolution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Tianyu, Lu, Shuai, Lu, Shan, Gong, Yeyun, Yang, Chenyuan, Li, Xuheng, Misu, Md Rakib Hossain, Yu, Hao, Duan, Nan, Cheng, Peng, Yang, Fan, Lahiri, Shuvendu K, Xie, Tao, Zhou, Lidong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908833896464384
author Chen, Tianyu
Lu, Shuai
Lu, Shan
Gong, Yeyun
Yang, Chenyuan
Li, Xuheng
Misu, Md Rakib Hossain
Yu, Hao
Duan, Nan
Cheng, Peng
Yang, Fan
Lahiri, Shuvendu K
Xie, Tao
Zhou, Lidong
author_facet Chen, Tianyu
Lu, Shuai
Lu, Shan
Gong, Yeyun
Yang, Chenyuan
Li, Xuheng
Misu, Md Rakib Hossain
Yu, Hao
Duan, Nan
Cheng, Peng
Yang, Fan
Lahiri, Shuvendu K
Xie, Tao
Zhou, Lidong
contents Ensuring correctness is crucial for code generation. Formal verification offers a definitive assurance of correctness, but demands substantial human effort in proof construction and hence raises a pressing need for automation. The primary obstacle lies in the severe lack of data-there is much fewer proofs than code snippets for Large Language Models (LLMs) to train upon. In this paper, we introduce SAFE, a framework that overcomes the lack of human-written proofs to enable automated proof generation of Rust code. SAFE establishes a self-evolving cycle where data synthesis and fine-tuning collaborate to enhance the model capability, leveraging the definitive power of a symbolic verifier in telling correct proofs from incorrect ones. SAFE also re-purposes the large number of synthesized incorrect proofs to train the self-debugging capability of the fine-tuned models, empowering them to fix incorrect proofs based on the verifier's feedback. SAFE demonstrates superior efficiency and precision compared to GPT-4o. Through tens of thousands of synthesized proofs and the self-debugging mechanism, we improve the capability of open-source models, initially unacquainted with formal verification, to automatically write proofs for Rust code. This advancement leads to a significant improvement in performance, achieving a 52.52% accuracy rate in a benchmark crafted by human experts, a significant leap over GPT-4o's performance of 14.39%.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15756
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Automated Proof Generation for Rust Code via Self-Evolution
Chen, Tianyu
Lu, Shuai
Lu, Shan
Gong, Yeyun
Yang, Chenyuan
Li, Xuheng
Misu, Md Rakib Hossain
Yu, Hao
Duan, Nan
Cheng, Peng
Yang, Fan
Lahiri, Shuvendu K
Xie, Tao
Zhou, Lidong
Software Engineering
Artificial Intelligence
Ensuring correctness is crucial for code generation. Formal verification offers a definitive assurance of correctness, but demands substantial human effort in proof construction and hence raises a pressing need for automation. The primary obstacle lies in the severe lack of data-there is much fewer proofs than code snippets for Large Language Models (LLMs) to train upon. In this paper, we introduce SAFE, a framework that overcomes the lack of human-written proofs to enable automated proof generation of Rust code. SAFE establishes a self-evolving cycle where data synthesis and fine-tuning collaborate to enhance the model capability, leveraging the definitive power of a symbolic verifier in telling correct proofs from incorrect ones. SAFE also re-purposes the large number of synthesized incorrect proofs to train the self-debugging capability of the fine-tuned models, empowering them to fix incorrect proofs based on the verifier's feedback. SAFE demonstrates superior efficiency and precision compared to GPT-4o. Through tens of thousands of synthesized proofs and the self-debugging mechanism, we improve the capability of open-source models, initially unacquainted with formal verification, to automatically write proofs for Rust code. This advancement leads to a significant improvement in performance, achieving a 52.52% accuracy rate in a benchmark crafted by human experts, a significant leap over GPT-4o's performance of 14.39%.
title Automated Proof Generation for Rust Code via Self-Evolution
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2410.15756