NL in the Middle: Code Translation with LLMs and Intermediate Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tai, Chi-en Amy, Nie, Pengyu, Golab, Lukasz, Wong, Alexander
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911158820143104
author Tai, Chi-en Amy
Nie, Pengyu
Golab, Lukasz
Wong, Alexander
author_facet Tai, Chi-en Amy
Nie, Pengyu
Golab, Lukasz
Wong, Alexander
contents Studies show that large language models (LLMs) produce buggy code translations. One promising avenue to improve translation accuracy is through intermediate representations, which provide structured guidance for the translation process. We investigate whether LLM-based code translation can benefit from intermediate representations, specifically in the form of natural language (NL) summaries and abstract syntax trees (ASTs). Since prompt engineering greatly affects LLM performance, we consider several ways to integrate these representations, from one-shot to chain-of-thought (CoT) prompting. Using Open GPT4 8X7B and specialized StarCoder and CodeGen models on popular code translation benchmarks (CodeNet and AVATAR), we find that CoT with an intermediate NL summary performs best, with an increase of 13.8% and 6.7%, respectively, in successful translations for the best-performing model (Open GPT4 8X7B) compared to the zero-shot prompt.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08627
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NL in the Middle: Code Translation with LLMs and Intermediate Representations
Tai, Chi-en Amy
Nie, Pengyu
Golab, Lukasz
Wong, Alexander
Software Engineering
Studies show that large language models (LLMs) produce buggy code translations. One promising avenue to improve translation accuracy is through intermediate representations, which provide structured guidance for the translation process. We investigate whether LLM-based code translation can benefit from intermediate representations, specifically in the form of natural language (NL) summaries and abstract syntax trees (ASTs). Since prompt engineering greatly affects LLM performance, we consider several ways to integrate these representations, from one-shot to chain-of-thought (CoT) prompting. Using Open GPT4 8X7B and specialized StarCoder and CodeGen models on popular code translation benchmarks (CodeNet and AVATAR), we find that CoT with an intermediate NL summary performs best, with an increase of 13.8% and 6.7%, respectively, in successful translations for the best-performing model (Open GPT4 8X7B) compared to the zero-shot prompt.
title NL in the Middle: Code Translation with LLMs and Intermediate Representations
topic Software Engineering
url https://arxiv.org/abs/2507.08627