Are Decoder-Only Large Language Models the Silver Bullet for Code Search?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Yuxuan, Liu, Mingwei, Ou, Guangsheng, Li, Anji, Dai, Dekun, Wang, Yanlin, Zheng, Zibin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917426709397504
author Chen, Yuxuan
Liu, Mingwei
Ou, Guangsheng
Li, Anji
Dai, Dekun
Wang, Yanlin
Zheng, Zibin
author_facet Chen, Yuxuan
Liu, Mingwei
Ou, Guangsheng
Li, Anji
Dai, Dekun
Wang, Yanlin
Zheng, Zibin
contents Code search is essential for code reuse, allowing developers to efficiently locate relevant code snippets. The advent of powerful decoder-only Large Language Models (LLMs) has revolutionized many code intelligence tasks. However, their effectiveness for the retrieval-based task of code search, particularly compared to established encoder-based models, remains underexplored. This paper addresses this gap by presenting a large-scale systematic evaluation of eleven decoder-only LLMs, analyzing their performance across zero-shot and fine-tuned settings. Our results show that fine-tuned decoder-only models, particularly CodeGemma, significantly outperform encoder-only models like UniXcoder, achieving a 40.4% higher Mean Average Precision (MAP) on the CoSQA$^+$ benchmark. Our analysis further reveals two crucial nuances for practitioners: first, the relationship between model size and performance is non-monotonic, with mid-sized models often outperforming larger variants; second, the composition of the training data is critical, as a multilingual dataset enhances generalization while a small amount of data from a specific language can act as noise and interfere with model effectiveness. These findings offer a comprehensive guide to selecting and optimizing modern LLMs for code search.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22240
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Are Decoder-Only Large Language Models the Silver Bullet for Code Search?
Chen, Yuxuan
Liu, Mingwei
Ou, Guangsheng
Li, Anji
Dai, Dekun
Wang, Yanlin
Zheng, Zibin
Software Engineering
D.2
Code search is essential for code reuse, allowing developers to efficiently locate relevant code snippets. The advent of powerful decoder-only Large Language Models (LLMs) has revolutionized many code intelligence tasks. However, their effectiveness for the retrieval-based task of code search, particularly compared to established encoder-based models, remains underexplored. This paper addresses this gap by presenting a large-scale systematic evaluation of eleven decoder-only LLMs, analyzing their performance across zero-shot and fine-tuned settings. Our results show that fine-tuned decoder-only models, particularly CodeGemma, significantly outperform encoder-only models like UniXcoder, achieving a 40.4% higher Mean Average Precision (MAP) on the CoSQA$^+$ benchmark. Our analysis further reveals two crucial nuances for practitioners: first, the relationship between model size and performance is non-monotonic, with mid-sized models often outperforming larger variants; second, the composition of the training data is critical, as a multilingual dataset enhances generalization while a small amount of data from a specific language can act as noise and interfere with model effectiveness. These findings offer a comprehensive guide to selecting and optimizing modern LLMs for code search.
title Are Decoder-Only Large Language Models the Silver Bullet for Code Search?
topic Software Engineering
D.2
url https://arxiv.org/abs/2410.22240