An Empirical Study on LLM-based Agents for Automated Bug Fixing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Xiangxin, Ma, Zexiong, Gao, Pengfei, Peng, Chao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909854832001024
author Meng, Xiangxin
Ma, Zexiong
Gao, Pengfei
Peng, Chao
author_facet Meng, Xiangxin
Ma, Zexiong
Gao, Pengfei
Peng, Chao
contents Large language models (LLMs) and LLM-based Agents have been applied to fix bugs automatically, demonstrating the capability in addressing software defects by engaging in development environment interaction, iterative validation and code modification. However, systematic analysis of these agent systems remain limited, particularly regarding performance variations among top-performing ones. In this paper, we examine six repair systems on the SWE-bench Verified benchmark for automated bug fixing. We first assess each system's overall performance, noting the instances solvable by all or none of these systems, and explore the capabilities of different systems. We also compare fault localization accuracy at file and code symbol levels and evaluate bug reproduction capabilities. Through analysis, we concluded that further optimization is needed in both the LLM capability itself and the design of Agentic flow to improve the effectiveness of the Agent in bug fixing.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10213
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Empirical Study on LLM-based Agents for Automated Bug Fixing
Meng, Xiangxin
Ma, Zexiong
Gao, Pengfei
Peng, Chao
Software Engineering
Artificial Intelligence
Large language models (LLMs) and LLM-based Agents have been applied to fix bugs automatically, demonstrating the capability in addressing software defects by engaging in development environment interaction, iterative validation and code modification. However, systematic analysis of these agent systems remain limited, particularly regarding performance variations among top-performing ones. In this paper, we examine six repair systems on the SWE-bench Verified benchmark for automated bug fixing. We first assess each system's overall performance, noting the instances solvable by all or none of these systems, and explore the capabilities of different systems. We also compare fault localization accuracy at file and code symbol levels and evaluate bug reproduction capabilities. Through analysis, we concluded that further optimization is needed in both the LLM capability itself and the design of Agentic flow to improve the effectiveness of the Agent in bug fixing.
title An Empirical Study on LLM-based Agents for Automated Bug Fixing
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2411.10213