Beyond Translation Accuracy: Addressing False Failures in LLM-Based Code Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rabbi, Fazle, Saha, Soumit Kanti, Yang, Jinqiu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917470846058496
author Rabbi, Fazle
Saha, Soumit Kanti
Yang, Jinqiu
author_facet Rabbi, Fazle
Saha, Soumit Kanti
Yang, Jinqiu
contents Large Language Models (LLMs) have achieved remarkable success in automated code translation. While prior work has focused on improving translation accuracy through advanced prompting and iterative repair, the reliability of the underlying evaluation frameworks has received less attention. In this paper, we demonstrate that a significant number of reported failures in code translation are not due to incorrect logic, but rather evaluation-induced errors stemming from improper compilation flags, missing library links, and unconfigured runtime environments. We conduct a large-scale empirical study across five programming languages (C, C++, Java, Python, Go) and three benchmarks (Avatar, CodeNet, EvalPlus), covering 6,164 translations generated by GPT-4o, DeepSeek-Coder, and Magicoder. Our analysis identifies and categorizes common false negatives, distinguishing pipeline-induced failures that affect any model from model-dependent behaviors that vary across LLMs. Our findings highlight the necessity for transparent, configuration-aware evaluation standards to accurately assess progress in LLM-based code translation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02195
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Translation Accuracy: Addressing False Failures in LLM-Based Code Translation
Rabbi, Fazle
Saha, Soumit Kanti
Yang, Jinqiu
Software Engineering
Large Language Models (LLMs) have achieved remarkable success in automated code translation. While prior work has focused on improving translation accuracy through advanced prompting and iterative repair, the reliability of the underlying evaluation frameworks has received less attention. In this paper, we demonstrate that a significant number of reported failures in code translation are not due to incorrect logic, but rather evaluation-induced errors stemming from improper compilation flags, missing library links, and unconfigured runtime environments. We conduct a large-scale empirical study across five programming languages (C, C++, Java, Python, Go) and three benchmarks (Avatar, CodeNet, EvalPlus), covering 6,164 translations generated by GPT-4o, DeepSeek-Coder, and Magicoder. Our analysis identifies and categorizes common false negatives, distinguishing pipeline-induced failures that affect any model from model-dependent behaviors that vary across LLMs. Our findings highlight the necessity for transparent, configuration-aware evaluation standards to accurately assess progress in LLM-based code translation.
title Beyond Translation Accuracy: Addressing False Failures in LLM-Based Code Translation
topic Software Engineering
url https://arxiv.org/abs/2605.02195