CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Ye, Meng, Rui, Joty, Shafiq, Savarese, Silvio, Xiong, Caiming, Zhou, Yingbo, Yavuz, Semih
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918117859393536
author Liu, Ye
Meng, Rui
Joty, Shafiq
Savarese, Silvio
Xiong, Caiming
Zhou, Yingbo
Yavuz, Semih
author_facet Liu, Ye
Meng, Rui
Joty, Shafiq
Savarese, Silvio
Xiong, Caiming
Zhou, Yingbo
Yavuz, Semih
contents Despite the success of text retrieval in many NLP tasks, code retrieval remains a largely underexplored area. Most text retrieval systems are tailored for natural language queries, often neglecting the specific challenges of retrieving code. This gap leaves existing models unable to effectively capture the diversity of programming languages and tasks across different domains, highlighting the need for more focused research in code retrieval. To address this, we introduce CodeXEmbed, a family of large-scale code embedding models ranging from 400M to 7B parameters. Our novel training pipeline unifies multiple programming languages and transforms various code-related tasks into a common retrieval framework, enhancing model generalizability and retrieval performance. Our 7B model sets a new state-of-the-art (SOTA) in code retrieval, outperforming the previous leading model, Voyage-Code, by over 20% on CoIR benchmark. In addition to excelling in code retrieval, our models demonstrate competitive performance on the widely adopted BeIR text retrieval benchmark, offering versatility across domains. Experimental results demonstrate that improving retrieval performance significantly enhances end-to-end Retrieval-Augmented Generation (RAG) performance for code-related tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12644
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
Liu, Ye
Meng, Rui
Joty, Shafiq
Savarese, Silvio
Xiong, Caiming
Zhou, Yingbo
Yavuz, Semih
Software Engineering
Artificial Intelligence
Despite the success of text retrieval in many NLP tasks, code retrieval remains a largely underexplored area. Most text retrieval systems are tailored for natural language queries, often neglecting the specific challenges of retrieving code. This gap leaves existing models unable to effectively capture the diversity of programming languages and tasks across different domains, highlighting the need for more focused research in code retrieval. To address this, we introduce CodeXEmbed, a family of large-scale code embedding models ranging from 400M to 7B parameters. Our novel training pipeline unifies multiple programming languages and transforms various code-related tasks into a common retrieval framework, enhancing model generalizability and retrieval performance. Our 7B model sets a new state-of-the-art (SOTA) in code retrieval, outperforming the previous leading model, Voyage-Code, by over 20% on CoIR benchmark. In addition to excelling in code retrieval, our models demonstrate competitive performance on the widely adopted BeIR text retrieval benchmark, offering versatility across domains. Experimental results demonstrate that improving retrieval performance significantly enhances end-to-end Retrieval-Augmented Generation (RAG) performance for code-related tasks.
title CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2411.12644