ExeCoder: Empowering Large Language Models with Executability Representation for Code Translation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: He, Minghua, Chen, Yue, Yang, Fangkai, Zhao, Pu, Yin, Wenjie, Kang, Yu, Lin, Qingwei, Rajmohan, Saravan, Zhang, Dongmei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916973792722944
author He, Minghua
Chen, Yue
Yang, Fangkai
Zhao, Pu
Yin, Wenjie
Kang, Yu
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
author_facet He, Minghua
Chen, Yue
Yang, Fangkai
Zhao, Pu
Yin, Wenjie
Kang, Yu
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
contents Code translation is a crucial activity in the software development and maintenance process, and researchers have recently begun to focus on using pre-trained large language models (LLMs) for code translation. However, existing LLMs only learn the contextual semantics of code during pre-training, neglecting executability information closely related to the execution state of the code, which results in unguaranteed code executability and unreliable automated code translation. To address this issue, we propose ExeCoder, an LLM specifically designed for code translation, aimed at utilizing executability representations such as functional semantics, syntax structures, and variable dependencies to enhance the capabilities of LLMs in code translation. To evaluate the effectiveness of ExeCoder, we manually enhanced the widely used benchmark TransCoder-test, resulting in a benchmark called TransCoder-test-X that serves LLMs. Evaluation of TransCoder-test-X indicates that ExeCoder achieves state-of-the-art performance in code translation, surpassing existing open-source code LLMs by over 10.88% to 38.78% and over 27.44% to 42.97% on two metrics, and even outperforms the renowned closed-source LLM GPT-4o. Code is available at https://aka.ms/execoder
format Preprint
id arxiv_https___arxiv_org_abs_2501_18460
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ExeCoder: Empowering Large Language Models with Executability Representation for Code Translation
He, Minghua
Chen, Yue
Yang, Fangkai
Zhao, Pu
Yin, Wenjie
Kang, Yu
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
Software Engineering
Code translation is a crucial activity in the software development and maintenance process, and researchers have recently begun to focus on using pre-trained large language models (LLMs) for code translation. However, existing LLMs only learn the contextual semantics of code during pre-training, neglecting executability information closely related to the execution state of the code, which results in unguaranteed code executability and unreliable automated code translation. To address this issue, we propose ExeCoder, an LLM specifically designed for code translation, aimed at utilizing executability representations such as functional semantics, syntax structures, and variable dependencies to enhance the capabilities of LLMs in code translation. To evaluate the effectiveness of ExeCoder, we manually enhanced the widely used benchmark TransCoder-test, resulting in a benchmark called TransCoder-test-X that serves LLMs. Evaluation of TransCoder-test-X indicates that ExeCoder achieves state-of-the-art performance in code translation, surpassing existing open-source code LLMs by over 10.88% to 38.78% and over 27.44% to 42.97% on two metrics, and even outperforms the renowned closed-source LLM GPT-4o. Code is available at https://aka.ms/execoder
title ExeCoder: Empowering Large Language Models with Executability Representation for Code Translation
topic Software Engineering
url https://arxiv.org/abs/2501.18460