Extracting Conceptual Knowledge to Locate Software Issues

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ying, Mao, Wenjun, Wang, Chong, Zhou, Zhenhao, Zhou, Yicheng, Zhao, Wenyun, Lou, Yiling, Peng, Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914074882736128
author Wang, Ying
Mao, Wenjun
Wang, Chong
Zhou, Zhenhao
Zhou, Yicheng
Zhao, Wenyun
Lou, Yiling
Peng, Xin
author_facet Wang, Ying
Mao, Wenjun
Wang, Chong
Zhou, Zhenhao
Zhou, Yicheng
Zhao, Wenyun
Lou, Yiling
Peng, Xin
contents Issue localization, which identifies faulty code elements such as files or functions, is critical for effective bug fixing. While recent LLM-based and LLM-agent-based approaches improve accuracy, they struggle in large-scale repositories due to concern tangling, where relevant logic is buried in large functions, and concern scattering, where related logic is dispersed across files. To address these challenges, we propose RepoLens, a novel approach that abstracts and leverages conceptual knowledge from code repositories. RepoLens decomposes fine-grained functionalities and recomposes them into high-level concerns, semantically coherent clusters of functionalities that guide LLMs. It operates in two stages: an offline stage that extracts and enriches conceptual knowledge into a repository-wide knowledge base, and an online stage that retrieves issue-specific terms, clusters and ranks concerns by relevance, and integrates them into localization workflows via minimally intrusive prompt enhancements. We evaluate RepoLens on SWE-Lancer-Loc, a benchmark of 216 tasks derived from SWE-Lancer. RepoLens consistently improves three state-of-the-art tools, namely AgentLess, OpenHands, and mini-SWE-agent, achieving average gains of over 22% in Hit@k and 46% in Recall@k for file- and function-level localization. It generalizes across models (GPT-4o, GPT-4o-mini, GPT-4.1) with Hit@1 and Recall@10 gains up to 504% and 376%, respectively. Ablation studies and manual evaluation confirm the effectiveness and reliability of the constructed concerns.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21427
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Extracting Conceptual Knowledge to Locate Software Issues
Wang, Ying
Mao, Wenjun
Wang, Chong
Zhou, Zhenhao
Zhou, Yicheng
Zhao, Wenyun
Lou, Yiling
Peng, Xin
Software Engineering
Issue localization, which identifies faulty code elements such as files or functions, is critical for effective bug fixing. While recent LLM-based and LLM-agent-based approaches improve accuracy, they struggle in large-scale repositories due to concern tangling, where relevant logic is buried in large functions, and concern scattering, where related logic is dispersed across files. To address these challenges, we propose RepoLens, a novel approach that abstracts and leverages conceptual knowledge from code repositories. RepoLens decomposes fine-grained functionalities and recomposes them into high-level concerns, semantically coherent clusters of functionalities that guide LLMs. It operates in two stages: an offline stage that extracts and enriches conceptual knowledge into a repository-wide knowledge base, and an online stage that retrieves issue-specific terms, clusters and ranks concerns by relevance, and integrates them into localization workflows via minimally intrusive prompt enhancements. We evaluate RepoLens on SWE-Lancer-Loc, a benchmark of 216 tasks derived from SWE-Lancer. RepoLens consistently improves three state-of-the-art tools, namely AgentLess, OpenHands, and mini-SWE-agent, achieving average gains of over 22% in Hit@k and 46% in Recall@k for file- and function-level localization. It generalizes across models (GPT-4o, GPT-4o-mini, GPT-4.1) with Hit@1 and Recall@10 gains up to 504% and 376%, respectively. Ablation studies and manual evaluation confirm the effectiveness and reliability of the constructed concerns.
title Extracting Conceptual Knowledge to Locate Software Issues
topic Software Engineering
url https://arxiv.org/abs/2509.21427