A Benchmark for Localizing Code and Non-Code Issues in Software Projects

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Zejun, Wang, Jian, Yang, Qingyun, Pan, Yifan, Tang, Yi, Li, Yi, Xing, Zhenchang, Zhang, Tian, Li, Xuandong, Zhang, Guoan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911184016375808
author Zhang, Zejun
Wang, Jian
Yang, Qingyun
Pan, Yifan
Tang, Yi
Li, Yi
Xing, Zhenchang
Zhang, Tian
Li, Xuandong
Zhang, Guoan
author_facet Zhang, Zejun
Wang, Jian
Yang, Qingyun
Pan, Yifan
Tang, Yi
Li, Yi
Xing, Zhenchang
Zhang, Tian
Li, Xuandong
Zhang, Guoan
contents Accurate project localization (e.g., files and functions) for issue resolution is a critical first step in software maintenance. However, existing benchmarks for issue localization, such as SWE-Bench and LocBench, are limited. They focus predominantly on pull-request issues and code locations, ignoring other evidence and non-code files such as commits, comments, configurations, and documentation. To address this gap, we introduce MULocBench, a comprehensive dataset of 1,100 issues from 46 popular GitHub Python projects. Comparing with existing benchmarks, MULocBench offers greater diversity in issue types, root causes, location scopes, and file types, providing a more realistic testbed for evaluation. Using this benchmark, we assess the performance of state-of-the-art localization methods and five LLM-based prompting strategies. Our results reveal significant limitations in current techniques: even at the file level, performance metrics (Acc@5, F1) remain below 40%. This underscores the challenge of generalizing to realistic, multi-faceted issue resolution. To enable future research on project localization for issue resolution, we publicly release MULocBench at https://huggingface.co/datasets/somethingone/MULocBench.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25242
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Benchmark for Localizing Code and Non-Code Issues in Software Projects
Zhang, Zejun
Wang, Jian
Yang, Qingyun
Pan, Yifan
Tang, Yi
Li, Yi
Xing, Zhenchang
Zhang, Tian
Li, Xuandong
Zhang, Guoan
Software Engineering
Artificial Intelligence
Accurate project localization (e.g., files and functions) for issue resolution is a critical first step in software maintenance. However, existing benchmarks for issue localization, such as SWE-Bench and LocBench, are limited. They focus predominantly on pull-request issues and code locations, ignoring other evidence and non-code files such as commits, comments, configurations, and documentation. To address this gap, we introduce MULocBench, a comprehensive dataset of 1,100 issues from 46 popular GitHub Python projects. Comparing with existing benchmarks, MULocBench offers greater diversity in issue types, root causes, location scopes, and file types, providing a more realistic testbed for evaluation. Using this benchmark, we assess the performance of state-of-the-art localization methods and five LLM-based prompting strategies. Our results reveal significant limitations in current techniques: even at the file level, performance metrics (Acc@5, F1) remain below 40%. This underscores the challenge of generalizing to realistic, multi-faceted issue resolution. To enable future research on project localization for issue resolution, we publicly release MULocBench at https://huggingface.co/datasets/somethingone/MULocBench.
title A Benchmark for Localizing Code and Non-Code Issues in Software Projects
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2509.25242