DebugBench: Evaluating Debugging Capability of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Runchu, Ye, Yining, Qin, Yujia, Cong, Xin, Lin, Yankai, Pan, Yinxu, Wu, Yesai, Hui, Haotian, Liu, Weichuan, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929375700582400
author Tian, Runchu
Ye, Yining
Qin, Yujia
Cong, Xin
Lin, Yankai
Pan, Yinxu
Wu, Yesai
Hui, Haotian
Liu, Weichuan
Liu, Zhiyuan
Sun, Maosong
author_facet Tian, Runchu
Ye, Yining
Qin, Yujia
Cong, Xin
Lin, Yankai
Pan, Yinxu
Wu, Yesai
Hui, Haotian
Liu, Weichuan
Liu, Zhiyuan
Sun, Maosong
contents Large Language Models (LLMs) have demonstrated exceptional coding capability. However, as another critical component of programming proficiency, the debugging capability of LLMs remains relatively unexplored. Previous evaluations of LLMs' debugging ability are significantly limited by the risk of data leakage, the scale of the dataset, and the variety of tested bugs. To overcome these deficiencies, we introduce `DebugBench', an LLM debugging benchmark consisting of 4,253 instances. It covers four major bug categories and 18 minor types in C++, Java, and Python. To construct DebugBench, we collect code snippets from the LeetCode community, implant bugs into source data with GPT-4, and assure rigorous quality checks. We evaluate two commercial and four open-source models in a zero-shot scenario. We find that (1) while closed-source models exhibit inferior debugging performance compared to humans, open-source models relatively lower pass rate scores; (2) the complexity of debugging notably fluctuates depending on the bug category; (3) incorporating runtime feedback has a clear impact on debugging performance which is not always helpful. As an extension, we also compare LLM debugging and code generation, revealing a strong correlation between them for closed-source models. These findings will benefit the development of LLMs in debugging.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04621
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DebugBench: Evaluating Debugging Capability of Large Language Models
Tian, Runchu
Ye, Yining
Qin, Yujia
Cong, Xin
Lin, Yankai
Pan, Yinxu
Wu, Yesai
Hui, Haotian
Liu, Weichuan
Liu, Zhiyuan
Sun, Maosong
Software Engineering
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) have demonstrated exceptional coding capability. However, as another critical component of programming proficiency, the debugging capability of LLMs remains relatively unexplored. Previous evaluations of LLMs' debugging ability are significantly limited by the risk of data leakage, the scale of the dataset, and the variety of tested bugs. To overcome these deficiencies, we introduce `DebugBench', an LLM debugging benchmark consisting of 4,253 instances. It covers four major bug categories and 18 minor types in C++, Java, and Python. To construct DebugBench, we collect code snippets from the LeetCode community, implant bugs into source data with GPT-4, and assure rigorous quality checks. We evaluate two commercial and four open-source models in a zero-shot scenario. We find that (1) while closed-source models exhibit inferior debugging performance compared to humans, open-source models relatively lower pass rate scores; (2) the complexity of debugging notably fluctuates depending on the bug category; (3) incorporating runtime feedback has a clear impact on debugging performance which is not always helpful. As an extension, we also compare LLM debugging and code generation, revealing a strong correlation between them for closed-source models. These findings will benefit the development of LLMs in debugging.
title DebugBench: Evaluating Debugging Capability of Large Language Models
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2401.04621