MdEval: Massively Multilingual Code Debugging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shukai, Chai, Linzheng, Yang, Jian, Shi, Jiajun, Zhu, He, Wang, Liran, Jin, Ke, Zhang, Wei, Zhu, Hualei, Guo, Shuyue, Sun, Tao, Liu, Jiaheng, Duan, Yunlong, Hao, Yu, Yang, Liqun, Niu, Guanglin, Zhang, Ge, Li, Zhoujun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917934858764288
author Liu, Shukai
Chai, Linzheng
Yang, Jian
Shi, Jiajun
Zhu, He
Wang, Liran
Jin, Ke
Zhang, Wei
Zhu, Hualei
Guo, Shuyue
Sun, Tao
Liu, Jiaheng
Duan, Yunlong
Hao, Yu
Yang, Liqun
Niu, Guanglin
Zhang, Ge
Li, Zhoujun
author_facet Liu, Shukai
Chai, Linzheng
Yang, Jian
Shi, Jiajun
Zhu, He
Wang, Liran
Jin, Ke
Zhang, Wei
Zhu, Hualei
Guo, Shuyue
Sun, Tao
Liu, Jiaheng
Duan, Yunlong
Hao, Yu
Yang, Liqun
Niu, Guanglin
Zhang, Ge
Li, Zhoujun
contents Code large language models (LLMs) have made significant progress in code debugging by directly generating the correct code based on the buggy code snippet. Programming benchmarks, typically consisting of buggy code snippet and their associated test cases, are used to assess the debugging capabilities of LLMs. However, many existing benchmarks primarily focus on Python and are often limited in terms of language diversity (e.g., DebugBench and DebugEval). To advance the field of multilingual debugging with LLMs, we propose the first massively multilingual debugging benchmark, which includes 3.6K test samples of 18 programming languages and covers the automated program repair (APR) task, the code review (CR) task, and the bug identification (BI) task. Further, we introduce the debugging instruction corpora MDEVAL-INSTRUCT by injecting bugs into the correct multilingual queries and solutions (xDebugGen). Further, a multilingual debugger xDebugCoder trained on MDEVAL-INSTRUCT as a strong baseline specifically to handle the bugs of a wide range of programming languages (e.g. "Missing Mut" in language Rust and "Misused Macro Definition" in language C). Our extensive experiments on MDEVAL reveal a notable performance gap between open-source models and closed-source LLMs (e.g., GPT and Claude series), highlighting huge room for improvement in multilingual code debugging scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02310
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MdEval: Massively Multilingual Code Debugging
Liu, Shukai
Chai, Linzheng
Yang, Jian
Shi, Jiajun
Zhu, He
Wang, Liran
Jin, Ke
Zhang, Wei
Zhu, Hualei
Guo, Shuyue
Sun, Tao
Liu, Jiaheng
Duan, Yunlong
Hao, Yu
Yang, Liqun
Niu, Guanglin
Zhang, Ge
Li, Zhoujun
Computation and Language
Code large language models (LLMs) have made significant progress in code debugging by directly generating the correct code based on the buggy code snippet. Programming benchmarks, typically consisting of buggy code snippet and their associated test cases, are used to assess the debugging capabilities of LLMs. However, many existing benchmarks primarily focus on Python and are often limited in terms of language diversity (e.g., DebugBench and DebugEval). To advance the field of multilingual debugging with LLMs, we propose the first massively multilingual debugging benchmark, which includes 3.6K test samples of 18 programming languages and covers the automated program repair (APR) task, the code review (CR) task, and the bug identification (BI) task. Further, we introduce the debugging instruction corpora MDEVAL-INSTRUCT by injecting bugs into the correct multilingual queries and solutions (xDebugGen). Further, a multilingual debugger xDebugCoder trained on MDEVAL-INSTRUCT as a strong baseline specifically to handle the bugs of a wide range of programming languages (e.g. "Missing Mut" in language Rust and "Misused Macro Definition" in language C). Our extensive experiments on MDEVAL reveal a notable performance gap between open-source models and closed-source LLMs (e.g., GPT and Claude series), highlighting huge room for improvement in multilingual code debugging scenarios.
title MdEval: Massively Multilingual Code Debugging
topic Computation and Language
url https://arxiv.org/abs/2411.02310