SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912002266365952 |
|---|---|
| author | Zan, Daoguang Huang, Zhirong Yu, Ailun Lin, Shaoxin Shi, Yifan Liu, Wei Chen, Dong Qi, Zongshuai Yu, Hao Yu, Lei Ran, Dezhi Zeng, Muhan Shen, Bo Bian, Pan Liang, Guangtai Guan, Bei Huang, Pengjie Xie, Tao Wang, Yongji Wang, Qianxiang |
| author_facet | Zan, Daoguang Huang, Zhirong Yu, Ailun Lin, Shaoxin Shi, Yifan Liu, Wei Chen, Dong Qi, Zongshuai Yu, Hao Yu, Lei Ran, Dezhi Zeng, Muhan Shen, Bo Bian, Pan Liang, Guangtai Guan, Bei Huang, Pengjie Xie, Tao Wang, Yongji Wang, Qianxiang |
| contents | GitHub issue resolving is a critical task in software engineering, recently gaining significant attention in both industry and academia. Within this task, SWE-bench has been released to evaluate issue resolving capabilities of large language models (LLMs), but has so far only focused on Python version. However, supporting more programming languages is also important, as there is a strong demand in industry. As a first step toward multilingual support, we have developed a Java version of SWE-bench, called SWE-bench-java. We have publicly released the dataset, along with the corresponding Docker-based evaluation environment and leaderboard, which will be continuously maintained and updated in the coming months. To verify the reliability of SWE-bench-java, we implement a classic method SWE-agent and test several powerful LLMs on it. As is well known, developing a high-quality multi-lingual benchmark is time-consuming and labor-intensive, so we welcome contributions through pull requests or collaboration to accelerate its iteration and refinement, paving the way for fully automated programming. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_14354 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SWE-bench-java: A GitHub Issue Resolving Benchmark for Java Zan, Daoguang Huang, Zhirong Yu, Ailun Lin, Shaoxin Shi, Yifan Liu, Wei Chen, Dong Qi, Zongshuai Yu, Hao Yu, Lei Ran, Dezhi Zeng, Muhan Shen, Bo Bian, Pan Liang, Guangtai Guan, Bei Huang, Pengjie Xie, Tao Wang, Yongji Wang, Qianxiang Software Engineering Artificial Intelligence Computation and Language GitHub issue resolving is a critical task in software engineering, recently gaining significant attention in both industry and academia. Within this task, SWE-bench has been released to evaluate issue resolving capabilities of large language models (LLMs), but has so far only focused on Python version. However, supporting more programming languages is also important, as there is a strong demand in industry. As a first step toward multilingual support, we have developed a Java version of SWE-bench, called SWE-bench-java. We have publicly released the dataset, along with the corresponding Docker-based evaluation environment and leaderboard, which will be continuously maintained and updated in the coming months. To verify the reliability of SWE-bench-java, we implement a classic method SWE-agent and test several powerful LLMs on it. As is well known, developing a high-quality multi-lingual benchmark is time-consuming and labor-intensive, so we welcome contributions through pull requests or collaboration to accelerate its iteration and refinement, paving the way for fully automated programming. |
| title | SWE-bench-java: A GitHub Issue Resolving Benchmark for Java |
| topic | Software Engineering Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2408.14354 |