TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gong, Zhihao, Sun, Zeyu, Huang, Dong, Liang, Qingyuan, Zhang, Jie M., Hao, Dan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918445138837504
author Gong, Zhihao
Sun, Zeyu
Huang, Dong
Liang, Qingyuan
Zhang, Jie M.
Hao, Dan
author_facet Gong, Zhihao
Sun, Zeyu
Huang, Dong
Liang, Qingyuan
Zhang, Jie M.
Hao, Dan
contents While Large Language Models (LLMs) have substantially improved the functional correctness of code translation, the critical dimension of \textit{execution efficiency} remains overlooked. We present \textbf{\textsc{trace}}, the first benchmark to explicitly assess efficiency in LLM-translated code. \textsc{trace} includes 1,000 efficiency-critical tasks across C++, Java, and Python, each augmented with stress tests that reveal efficiency degradations often overlooked by small-scale tests. Using \textsc{trace}, we conduct an extensive evaluation of 28 representative LLMs and highlight several key insights: 1) Correctness is not a reliable proxy for efficiency: the correctness leader \textit{Claude-4-think} achieves only mid-level time efficiency, outperformed by smaller open-source LLMs such as \textit{Qwen2.5-Coder-14B-Instruct}. 2) Inefficiency is both prevalent and patterned: 23.5\% of correct translations exhibit pronounced inefficiency, distributed across algorithmic faults (11.9\%), language construct mismatches (66.4\%), and resource mismanagement (21.7\%). 3) Inference-time prompt strategies bring only modest improvements, suggesting that current LLMs lack intrinsic efficiency awareness. Together, our results establish efficiency as an essential dimension of code translation and position \textsc{trace} as a principled foundation for efficiency-oriented evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16479
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation
Gong, Zhihao
Sun, Zeyu
Huang, Dong
Liang, Qingyuan
Zhang, Jie M.
Hao, Dan
Software Engineering
While Large Language Models (LLMs) have substantially improved the functional correctness of code translation, the critical dimension of \textit{execution efficiency} remains overlooked. We present \textbf{\textsc{trace}}, the first benchmark to explicitly assess efficiency in LLM-translated code. \textsc{trace} includes 1,000 efficiency-critical tasks across C++, Java, and Python, each augmented with stress tests that reveal efficiency degradations often overlooked by small-scale tests. Using \textsc{trace}, we conduct an extensive evaluation of 28 representative LLMs and highlight several key insights: 1) Correctness is not a reliable proxy for efficiency: the correctness leader \textit{Claude-4-think} achieves only mid-level time efficiency, outperformed by smaller open-source LLMs such as \textit{Qwen2.5-Coder-14B-Instruct}. 2) Inefficiency is both prevalent and patterned: 23.5\% of correct translations exhibit pronounced inefficiency, distributed across algorithmic faults (11.9\%), language construct mismatches (66.4\%), and resource mismanagement (21.7\%). 3) Inference-time prompt strategies bring only modest improvements, suggesting that current LLMs lack intrinsic efficiency awareness. Together, our results establish efficiency as an essential dimension of code translation and position \textsc{trace} as a principled foundation for efficiency-oriented evaluation.
title TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation
topic Software Engineering
url https://arxiv.org/abs/2603.16479