Meta-RAG on Large Codebases Using Code Summarization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tawosi, Vali, Alamir, Salwa, Liu, Xiaomo, Veloso, Manuela
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916885269839872
author Tawosi, Vali
Alamir, Salwa
Liu, Xiaomo
Veloso, Manuela
author_facet Tawosi, Vali
Alamir, Salwa
Liu, Xiaomo
Veloso, Manuela
contents Large Language Model (LLM) systems have been at the forefront of applied Artificial Intelligence (AI) research in a multitude of domains. One such domain is software development, where researchers have pushed the automation of a number of code tasks through LLM agents. Software development is a complex ecosystem, that stretches far beyond code implementation and well into the realm of code maintenance. In this paper, we propose a multi-agent system to localize bugs in large pre-existing codebases using information retrieval and LLMs. Our system introduces a novel Retrieval Augmented Generation (RAG) approach, Meta-RAG, where we utilize summaries to condense codebases by an average of 79.8\%, into a compact, structured, natural language representation. We then use an LLM agent to determine which parts of the codebase are critical for bug resolution, i.e. bug localization. We demonstrate the usefulness of Meta-RAG through evaluation with the SWE-bench Lite dataset. Meta-RAG scores 84.67 % and 53.0 % for file-level and function-level correct localization rates, respectively, achieving state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02611
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Meta-RAG on Large Codebases Using Code Summarization
Tawosi, Vali
Alamir, Salwa
Liu, Xiaomo
Veloso, Manuela
Software Engineering
Artificial Intelligence
Large Language Model (LLM) systems have been at the forefront of applied Artificial Intelligence (AI) research in a multitude of domains. One such domain is software development, where researchers have pushed the automation of a number of code tasks through LLM agents. Software development is a complex ecosystem, that stretches far beyond code implementation and well into the realm of code maintenance. In this paper, we propose a multi-agent system to localize bugs in large pre-existing codebases using information retrieval and LLMs. Our system introduces a novel Retrieval Augmented Generation (RAG) approach, Meta-RAG, where we utilize summaries to condense codebases by an average of 79.8\%, into a compact, structured, natural language representation. We then use an LLM agent to determine which parts of the codebase are critical for bug resolution, i.e. bug localization. We demonstrate the usefulness of Meta-RAG through evaluation with the SWE-bench Lite dataset. Meta-RAG scores 84.67 % and 53.0 % for file-level and function-level correct localization rates, respectively, achieving state-of-the-art performance.
title Meta-RAG on Large Codebases Using Code Summarization
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2508.02611