Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheshkov, Anton, Zadorozhny, Pavel, Levichev, Rodion, Maslov, Evgeny, Jaldin, Ronaldo Franco
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909337436291072
author Cheshkov, Anton
Zadorozhny, Pavel
Levichev, Rodion
Maslov, Evgeny
Jaldin, Ronaldo Franco
author_facet Cheshkov, Anton
Zadorozhny, Pavel
Levichev, Rodion
Maslov, Evgeny
Jaldin, Ronaldo Franco
contents Automatic program repair at project level may open yet to be seen opportunities in various fields of human activity. Since the SWE-Bench challenge was presented, we have seen numerous of solutions. Patch generation is a part of program repair, and test suite-based conversational patch generation has proven its effectiveness. However, the potential of conversational patch generation has not yet specifically estimated on SWE-Bench. This study reports experimental results aimed at evaluating the individual effectiveness of conversational patch generation on problems from SWE-Bench. The experiments show that a simple conversational pipeline based on LLaMA 3.1 70B can generate valid patches in 47\% of cases, which is comparable to the state-of-the-art in program repair on SWE-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04485
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench
Cheshkov, Anton
Zadorozhny, Pavel
Levichev, Rodion
Maslov, Evgeny
Jaldin, Ronaldo Franco
Software Engineering
Artificial Intelligence
Multiagent Systems
Automatic program repair at project level may open yet to be seen opportunities in various fields of human activity. Since the SWE-Bench challenge was presented, we have seen numerous of solutions. Patch generation is a part of program repair, and test suite-based conversational patch generation has proven its effectiveness. However, the potential of conversational patch generation has not yet specifically estimated on SWE-Bench. This study reports experimental results aimed at evaluating the individual effectiveness of conversational patch generation on problems from SWE-Bench. The experiments show that a simple conversational pipeline based on LLaMA 3.1 70B can generate valid patches in 47\% of cases, which is comparable to the state-of-the-art in program repair on SWE-Bench.
title Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench
topic Software Engineering
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2410.04485