CoRet: Improved Retriever for Code Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fehr, Fabio, Sivaprasad, Prabhu Teja, Franceschi, Luca, Zappella, Giovanni
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909629466804224
author Fehr, Fabio
Sivaprasad, Prabhu Teja
Franceschi, Luca
Zappella, Giovanni
author_facet Fehr, Fabio
Sivaprasad, Prabhu Teja
Franceschi, Luca
Zappella, Giovanni
contents In this paper, we introduce CoRet, a dense retrieval model designed for code-editing tasks that integrates code semantics, repository structure, and call graph dependencies. The model focuses on retrieving relevant portions of a code repository based on natural language queries such as requests to implement new features or fix bugs. These retrieved code chunks can then be presented to a user or to a second code-editing model or agent. To train CoRet, we propose a loss function explicitly designed for repository-level retrieval. On SWE-bench and Long Code Arena's bug localisation datasets, we show that our model substantially improves retrieval recall by at least 15 percentage points over existing models, and ablate the design choices to show their importance in achieving these results.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24715
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoRet: Improved Retriever for Code Editing
Fehr, Fabio
Sivaprasad, Prabhu Teja
Franceschi, Luca
Zappella, Giovanni
Machine Learning
Artificial Intelligence
Computation and Language
In this paper, we introduce CoRet, a dense retrieval model designed for code-editing tasks that integrates code semantics, repository structure, and call graph dependencies. The model focuses on retrieving relevant portions of a code repository based on natural language queries such as requests to implement new features or fix bugs. These retrieved code chunks can then be presented to a user or to a second code-editing model or agent. To train CoRet, we propose a loss function explicitly designed for repository-level retrieval. On SWE-bench and Long Code Arena's bug localisation datasets, we show that our model substantially improves retrieval recall by at least 15 percentage points over existing models, and ablate the design choices to show their importance in achieving these results.
title CoRet: Improved Retriever for Code Editing
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.24715