SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zan, Daoguang, Huang, Zhirong, Yu, Ailun, Lin, Shaoxin, Shi, Yifan, Liu, Wei, Chen, Dong, Qi, Zongshuai, Yu, Hao, Yu, Lei, Ran, Dezhi, Zeng, Muhan, Shen, Bo, Bian, Pan, Liang, Guangtai, Guan, Bei, Huang, Pengjie, Xie, Tao, Wang, Yongji, Wang, Qianxiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912002266365952
author Zan, Daoguang
Huang, Zhirong
Yu, Ailun
Lin, Shaoxin
Shi, Yifan
Liu, Wei
Chen, Dong
Qi, Zongshuai
Yu, Hao
Yu, Lei
Ran, Dezhi
Zeng, Muhan
Shen, Bo
Bian, Pan
Liang, Guangtai
Guan, Bei
Huang, Pengjie
Xie, Tao
Wang, Yongji
Wang, Qianxiang
author_facet Zan, Daoguang
Huang, Zhirong
Yu, Ailun
Lin, Shaoxin
Shi, Yifan
Liu, Wei
Chen, Dong
Qi, Zongshuai
Yu, Hao
Yu, Lei
Ran, Dezhi
Zeng, Muhan
Shen, Bo
Bian, Pan
Liang, Guangtai
Guan, Bei
Huang, Pengjie
Xie, Tao
Wang, Yongji
Wang, Qianxiang
contents GitHub issue resolving is a critical task in software engineering, recently gaining significant attention in both industry and academia. Within this task, SWE-bench has been released to evaluate issue resolving capabilities of large language models (LLMs), but has so far only focused on Python version. However, supporting more programming languages is also important, as there is a strong demand in industry. As a first step toward multilingual support, we have developed a Java version of SWE-bench, called SWE-bench-java. We have publicly released the dataset, along with the corresponding Docker-based evaluation environment and leaderboard, which will be continuously maintained and updated in the coming months. To verify the reliability of SWE-bench-java, we implement a classic method SWE-agent and test several powerful LLMs on it. As is well known, developing a high-quality multi-lingual benchmark is time-consuming and labor-intensive, so we welcome contributions through pull requests or collaboration to accelerate its iteration and refinement, paving the way for fully automated programming.
format Preprint
id arxiv_https___arxiv_org_abs_2408_14354
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
Zan, Daoguang
Huang, Zhirong
Yu, Ailun
Lin, Shaoxin
Shi, Yifan
Liu, Wei
Chen, Dong
Qi, Zongshuai
Yu, Hao
Yu, Lei
Ran, Dezhi
Zeng, Muhan
Shen, Bo
Bian, Pan
Liang, Guangtai
Guan, Bei
Huang, Pengjie
Xie, Tao
Wang, Yongji
Wang, Qianxiang
Software Engineering
Artificial Intelligence
Computation and Language
GitHub issue resolving is a critical task in software engineering, recently gaining significant attention in both industry and academia. Within this task, SWE-bench has been released to evaluate issue resolving capabilities of large language models (LLMs), but has so far only focused on Python version. However, supporting more programming languages is also important, as there is a strong demand in industry. As a first step toward multilingual support, we have developed a Java version of SWE-bench, called SWE-bench-java. We have publicly released the dataset, along with the corresponding Docker-based evaluation environment and leaderboard, which will be continuously maintained and updated in the coming months. To verify the reliability of SWE-bench-java, we implement a classic method SWE-agent and test several powerful LLMs on it. As is well known, developing a high-quality multi-lingual benchmark is time-consuming and labor-intensive, so we welcome contributions through pull requests or collaboration to accelerate its iteration and refinement, paving the way for fully automated programming.
title SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2408.14354