SweRank: Software Issue Localization with Code Ranking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Reddy, Revanth Gangi, Suresh, Tarun, Doo, JaeHyeok, Liu, Ye, Nguyen, Xuan Phi, Zhou, Yingbo, Yavuz, Semih, Xiong, Caiming, Ji, Heng, Joty, Shafiq
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913052384821248
author Reddy, Revanth Gangi
Suresh, Tarun
Doo, JaeHyeok
Liu, Ye
Nguyen, Xuan Phi
Zhou, Yingbo
Yavuz, Semih
Xiong, Caiming
Ji, Heng
Joty, Shafiq
author_facet Reddy, Revanth Gangi
Suresh, Tarun
Doo, JaeHyeok
Liu, Ye
Nguyen, Xuan Phi
Zhou, Yingbo
Yavuz, Semih
Xiong, Caiming
Ji, Heng
Joty, Shafiq
contents Software issue localization, the task of identifying the precise code locations (files, classes, or functions) relevant to a natural language issue description (e.g., bug report, feature request), is a critical yet time-consuming aspect of software development. While recent LLM-based agentic approaches demonstrate promise, they often incur significant latency and cost due to complex multi-step reasoning and relying on closed-source LLMs. Alternatively, traditional code ranking models, typically optimized for query-to-code or code-to-code retrieval, struggle with the verbose and failure-descriptive nature of issue localization queries. To bridge this gap, we introduce SweRank, an efficient and effective retrieve-and-rerank framework for software issue localization. To facilitate training, we construct SweLoc, a large-scale dataset curated from public GitHub repositories, featuring real-world issue descriptions paired with corresponding code modifications. Empirical results on SWE-Bench-Lite and LocBench show that SweRank achieves state-of-the-art performance, outperforming both prior ranking models and costly agent-based systems using closed-source LLMs like Claude-3.5. Further, we demonstrate SweLoc's utility in enhancing various existing retriever and reranker models for issue localization, establishing the dataset as a valuable resource for the community.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SweRank: Software Issue Localization with Code Ranking
Reddy, Revanth Gangi
Suresh, Tarun
Doo, JaeHyeok
Liu, Ye
Nguyen, Xuan Phi
Zhou, Yingbo
Yavuz, Semih
Xiong, Caiming
Ji, Heng
Joty, Shafiq
Software Engineering
Artificial Intelligence
Information Retrieval
Software issue localization, the task of identifying the precise code locations (files, classes, or functions) relevant to a natural language issue description (e.g., bug report, feature request), is a critical yet time-consuming aspect of software development. While recent LLM-based agentic approaches demonstrate promise, they often incur significant latency and cost due to complex multi-step reasoning and relying on closed-source LLMs. Alternatively, traditional code ranking models, typically optimized for query-to-code or code-to-code retrieval, struggle with the verbose and failure-descriptive nature of issue localization queries. To bridge this gap, we introduce SweRank, an efficient and effective retrieve-and-rerank framework for software issue localization. To facilitate training, we construct SweLoc, a large-scale dataset curated from public GitHub repositories, featuring real-world issue descriptions paired with corresponding code modifications. Empirical results on SWE-Bench-Lite and LocBench show that SweRank achieves state-of-the-art performance, outperforming both prior ranking models and costly agent-based systems using closed-source LLMs like Claude-3.5. Further, we demonstrate SweLoc's utility in enhancing various existing retriever and reranker models for issue localization, establishing the dataset as a valuable resource for the community.
title SweRank: Software Issue Localization with Code Ranking
topic Software Engineering
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2505.07849