A Benchmark for Localizing Code and Non-Code Issues in Software Projects
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911184016375808 |
|---|---|
| author | Zhang, Zejun Wang, Jian Yang, Qingyun Pan, Yifan Tang, Yi Li, Yi Xing, Zhenchang Zhang, Tian Li, Xuandong Zhang, Guoan |
| author_facet | Zhang, Zejun Wang, Jian Yang, Qingyun Pan, Yifan Tang, Yi Li, Yi Xing, Zhenchang Zhang, Tian Li, Xuandong Zhang, Guoan |
| contents | Accurate project localization (e.g., files and functions) for issue resolution is a critical first step in software maintenance. However, existing benchmarks for issue localization, such as SWE-Bench and LocBench, are limited. They focus predominantly on pull-request issues and code locations, ignoring other evidence and non-code files such as commits, comments, configurations, and documentation. To address this gap, we introduce MULocBench, a comprehensive dataset of 1,100 issues from 46 popular GitHub Python projects. Comparing with existing benchmarks, MULocBench offers greater diversity in issue types, root causes, location scopes, and file types, providing a more realistic testbed for evaluation. Using this benchmark, we assess the performance of state-of-the-art localization methods and five LLM-based prompting strategies. Our results reveal significant limitations in current techniques: even at the file level, performance metrics (Acc@5, F1) remain below 40%. This underscores the challenge of generalizing to realistic, multi-faceted issue resolution. To enable future research on project localization for issue resolution, we publicly release MULocBench at https://huggingface.co/datasets/somethingone/MULocBench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_25242 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Benchmark for Localizing Code and Non-Code Issues in Software Projects Zhang, Zejun Wang, Jian Yang, Qingyun Pan, Yifan Tang, Yi Li, Yi Xing, Zhenchang Zhang, Tian Li, Xuandong Zhang, Guoan Software Engineering Artificial Intelligence Accurate project localization (e.g., files and functions) for issue resolution is a critical first step in software maintenance. However, existing benchmarks for issue localization, such as SWE-Bench and LocBench, are limited. They focus predominantly on pull-request issues and code locations, ignoring other evidence and non-code files such as commits, comments, configurations, and documentation. To address this gap, we introduce MULocBench, a comprehensive dataset of 1,100 issues from 46 popular GitHub Python projects. Comparing with existing benchmarks, MULocBench offers greater diversity in issue types, root causes, location scopes, and file types, providing a more realistic testbed for evaluation. Using this benchmark, we assess the performance of state-of-the-art localization methods and five LLM-based prompting strategies. Our results reveal significant limitations in current techniques: even at the file level, performance metrics (Acc@5, F1) remain below 40%. This underscores the challenge of generalizing to realistic, multi-faceted issue resolution. To enable future research on project localization for issue resolution, we publicly release MULocBench at https://huggingface.co/datasets/somethingone/MULocBench. |
| title | A Benchmark for Localizing Code and Non-Code Issues in Software Projects |
| topic | Software Engineering Artificial Intelligence |
| url | https://arxiv.org/abs/2509.25242 |