RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yanli, Wang, Yanlin, Wang, Suiquan, Guo, Daya, Chen, Jiachi, Grundy, John, Liu, Xilin, Ma, Yuchi, Mao, Mingzhi, Zhang, Hongyu, Zheng, Zibin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911320812552192
author Wang, Yanli
Wang, Yanlin
Wang, Suiquan
Guo, Daya
Chen, Jiachi
Grundy, John
Liu, Xilin
Ma, Yuchi
Mao, Mingzhi
Zhang, Hongyu
Zheng, Zibin
author_facet Wang, Yanli
Wang, Yanlin
Wang, Suiquan
Guo, Daya
Chen, Jiachi
Grundy, John
Liu, Xilin
Ma, Yuchi
Mao, Mingzhi
Zhang, Hongyu
Zheng, Zibin
contents Repository-level code translation refers to translating an entire code repository from one programming language to another while preserving the functionality of the source repository. Many benchmarks have been proposed to evaluate the performance of such code translators. However, previous benchmarks mostly provide fine-grained samples, focusing at either code snippet, function, or file-level code translation. Such benchmarks do not accurately reflect real-world demands, where entire repositories often need to be translated, involving longer code length and more complex functionalities. To address this gap, we propose a new benchmark, named RepoTransBench, which is a real-world multilingual repository-level code translation benchmark featuring 1,897 real-world repository samples across 13 language pairs with automatically executable test suites. Besides, we introduce RepoTransAgent, a general agent framework to perform repository-level code translation. We evaluate both our benchmark's challenges and agent's effectiveness using several methods and backbone LLMs, revealing that repository-level translation remains challenging, where the best-performing method achieves only a 32.8% success rate. Furthermore, our analysis reveals that translation difficulty varies significantly by language pair direction, with dynamic-to-static language translation being much more challenging than the reverse direction (achieving below 10% vs. static-to-dynamic at 45-63%). Finally, we conduct a detailed error analysis and highlight current LLMs' deficiencies in repository-level code translation, which could provide a reference for further improvements. We provide the code and data at https://github.com/DeepSoftwareAnalytics/RepoTransBench.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17744
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
Wang, Yanli
Wang, Yanlin
Wang, Suiquan
Guo, Daya
Chen, Jiachi
Grundy, John
Liu, Xilin
Ma, Yuchi
Mao, Mingzhi
Zhang, Hongyu
Zheng, Zibin
Software Engineering
Artificial Intelligence
Computation and Language
Repository-level code translation refers to translating an entire code repository from one programming language to another while preserving the functionality of the source repository. Many benchmarks have been proposed to evaluate the performance of such code translators. However, previous benchmarks mostly provide fine-grained samples, focusing at either code snippet, function, or file-level code translation. Such benchmarks do not accurately reflect real-world demands, where entire repositories often need to be translated, involving longer code length and more complex functionalities. To address this gap, we propose a new benchmark, named RepoTransBench, which is a real-world multilingual repository-level code translation benchmark featuring 1,897 real-world repository samples across 13 language pairs with automatically executable test suites. Besides, we introduce RepoTransAgent, a general agent framework to perform repository-level code translation. We evaluate both our benchmark's challenges and agent's effectiveness using several methods and backbone LLMs, revealing that repository-level translation remains challenging, where the best-performing method achieves only a 32.8% success rate. Furthermore, our analysis reveals that translation difficulty varies significantly by language pair direction, with dynamic-to-static language translation being much more challenging than the reverse direction (achieving below 10% vs. static-to-dynamic at 45-63%). Finally, we conduct a detailed error analysis and highlight current LLMs' deficiencies in repository-level code translation, which could provide a reference for further improvements. We provide the code and data at https://github.com/DeepSoftwareAnalytics/RepoTransBench.
title RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.17744