Empirical Evaluation of Large Language Models in Automated Program Repair

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Jiajun, Li, Fengjie, Qi, Xinzhu, Zhang, Hongyu, Jiang, Jiajun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909650025185280
author Sun, Jiajun
Li, Fengjie
Qi, Xinzhu
Zhang, Hongyu
Jiang, Jiajun
author_facet Sun, Jiajun
Li, Fengjie
Qi, Xinzhu
Zhang, Hongyu
Jiang, Jiajun
contents The increasing prevalence of software bugs has made automated program repair (APR) a key research focus. Large language models (LLMs) offer new opportunities for APR, but existing studies mostly rely on smaller, earlier-generation models and Java benchmarks. The repair capabilities of modern, large-scale LLMs across diverse languages and scenarios remain underexplored. To address this, we conduct a comprehensive empirical study of four open-source LLMs, CodeLlama, LLaMA, StarCoder, and DeepSeek-Coder, spanning 7B to 33B parameters, diverse architectures, and purposes. We evaluate them across two bug scenarios (enterprise-grades and algorithmic), three languages (Java, C/C++, Python), and four prompting strategies, analyzing over 600K generated patches on six benchmarks. Key findings include: (1) model specialization (e.g., CodeLlama) can outperform larger general-purpose models (e.g., LLaMA); (2) repair performance does not scale linearly with model size; (3) correct patches often appear early in generation; and (4) prompts significantly affect results. These insights offer practical guidance for designing effective and efficient LLM-based APR systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13186
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Empirical Evaluation of Large Language Models in Automated Program Repair
Sun, Jiajun
Li, Fengjie
Qi, Xinzhu
Zhang, Hongyu
Jiang, Jiajun
Software Engineering
The increasing prevalence of software bugs has made automated program repair (APR) a key research focus. Large language models (LLMs) offer new opportunities for APR, but existing studies mostly rely on smaller, earlier-generation models and Java benchmarks. The repair capabilities of modern, large-scale LLMs across diverse languages and scenarios remain underexplored. To address this, we conduct a comprehensive empirical study of four open-source LLMs, CodeLlama, LLaMA, StarCoder, and DeepSeek-Coder, spanning 7B to 33B parameters, diverse architectures, and purposes. We evaluate them across two bug scenarios (enterprise-grades and algorithmic), three languages (Java, C/C++, Python), and four prompting strategies, analyzing over 600K generated patches on six benchmarks. Key findings include: (1) model specialization (e.g., CodeLlama) can outperform larger general-purpose models (e.g., LLaMA); (2) repair performance does not scale linearly with model size; (3) correct patches often appear early in generation; and (4) prompts significantly affect results. These insights offer practical guidance for designing effective and efficient LLM-based APR systems.
title Empirical Evaluation of Large Language Models in Automated Program Repair
topic Software Engineering
url https://arxiv.org/abs/2506.13186