M2rc-Eval: Massively Multilingual Repository-level Code Completion Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jiaheng, Deng, Ken, Liu, Congnan, Yang, Jian, Liu, Shukai, Zhu, He, Zhao, Peng, Chai, Linzheng, Wu, Yanan, Jin, Ke, Zhang, Ge, Wang, Zekun, Zhang, Guoan, Xiang, Bangyu, Su, Wenbo, Zheng, Bo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929563444969472
author Liu, Jiaheng
Deng, Ken
Liu, Congnan
Yang, Jian
Liu, Shukai
Zhu, He
Zhao, Peng
Chai, Linzheng
Wu, Yanan
Jin, Ke
Zhang, Ge
Wang, Zekun
Zhang, Guoan
Xiang, Bangyu
Su, Wenbo
Zheng, Bo
author_facet Liu, Jiaheng
Deng, Ken
Liu, Congnan
Yang, Jian
Liu, Shukai
Zhu, He
Zhao, Peng
Chai, Linzheng
Wu, Yanan
Jin, Ke
Zhang, Ge
Wang, Zekun
Zhang, Guoan
Xiang, Bangyu
Su, Wenbo
Zheng, Bo
contents Repository-level code completion has drawn great attention in software engineering, and several benchmark datasets have been introduced. However, existing repository-level code completion benchmarks usually focus on a limited number of languages (<5), which cannot evaluate the general code intelligence abilities across different languages for existing code Large Language Models (LLMs). Besides, the existing benchmarks usually report overall average scores of different languages, where the fine-grained abilities in different completion scenarios are ignored. Therefore, to facilitate the research of code LLMs in multilingual scenarios, we propose a massively multilingual repository-level code completion benchmark covering 18 programming languages (called M2RC-EVAL), and two types of fine-grained annotations (i.e., bucket-level and semantic-level) on different completion scenarios are provided, where we obtain these annotations based on the parsed abstract syntax tree. Moreover, we also curate a massively multilingual instruction corpora M2RC- INSTRUCT dataset to improve the repository-level code completion abilities of existing code LLMs. Comprehensive experimental results demonstrate the effectiveness of our M2RC-EVAL and M2RC-INSTRUCT.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21157
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle M2rc-Eval: Massively Multilingual Repository-level Code Completion Evaluation
Liu, Jiaheng
Deng, Ken
Liu, Congnan
Yang, Jian
Liu, Shukai
Zhu, He
Zhao, Peng
Chai, Linzheng
Wu, Yanan
Jin, Ke
Zhang, Ge
Wang, Zekun
Zhang, Guoan
Xiang, Bangyu
Su, Wenbo
Zheng, Bo
Computation and Language
Software Engineering
Repository-level code completion has drawn great attention in software engineering, and several benchmark datasets have been introduced. However, existing repository-level code completion benchmarks usually focus on a limited number of languages (<5), which cannot evaluate the general code intelligence abilities across different languages for existing code Large Language Models (LLMs). Besides, the existing benchmarks usually report overall average scores of different languages, where the fine-grained abilities in different completion scenarios are ignored. Therefore, to facilitate the research of code LLMs in multilingual scenarios, we propose a massively multilingual repository-level code completion benchmark covering 18 programming languages (called M2RC-EVAL), and two types of fine-grained annotations (i.e., bucket-level and semantic-level) on different completion scenarios are provided, where we obtain these annotations based on the parsed abstract syntax tree. Moreover, we also curate a massively multilingual instruction corpora M2RC- INSTRUCT dataset to improve the repository-level code completion abilities of existing code LLMs. Comprehensive experimental results demonstrate the effectiveness of our M2RC-EVAL and M2RC-INSTRUCT.
title M2rc-Eval: Massively Multilingual Repository-level Code Completion Evaluation
topic Computation and Language
Software Engineering
url https://arxiv.org/abs/2410.21157