Integrating Rules and Semantics for LLM-Based C-to-Rust Translation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Luo, Feng, Ji, Kexing, Gao, Cuiyun, Gao, Shuzheng, Feng, Jia, Liu, Kui, Xia, Xin, Lyu, Michael R.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911100471083008
author Luo, Feng
Ji, Kexing
Gao, Cuiyun
Gao, Shuzheng
Feng, Jia
Liu, Kui
Xia, Xin
Lyu, Michael R.
author_facet Luo, Feng
Ji, Kexing
Gao, Cuiyun
Gao, Shuzheng
Feng, Jia
Liu, Kui
Xia, Xin
Lyu, Michael R.
contents Automated translation of legacy C code into Rust aims to ensure memory safety while reducing the burden of manual migration. Early approaches in code translation rely on static rule-based methods, but they suffer from limited coverage due to dependence on predefined rule patterns. Recent works regard the task as a sequence-to-sequence problem by leveraging large language models (LLMs). Although these LLM-based methods are capable of reducing unsafe code blocks, the translated code often exhibits issues in following Rust rules and maintaining semantic consistency. On one hand, existing methods adopt a direct prompting strategy to translate the C code, which struggles to accommodate the syntactic rules between C and Rust. On the other hand, this strategy makes it difficult for LLMs to accurately capture the semantics of complex code. To address these challenges, we propose IRENE, an LLM-based framework that Integrates RulEs aNd sEmantics to enhance translation. IRENE consists of three modules: 1) a rule-augmented retrieval module that selects relevant translation examples based on rules generated from a static analyzer developed by us, thereby improving the handling of Rust rules; 2) a structured summarization module that produces a structured summary for guiding LLMs to enhance the semantic understanding of C code; 3) an error-driven translation module that leverages compiler diagnostics to iteratively refine translations. We evaluate IRENE on two datasets (xCodeEval, a public dataset, and HW-Bench, an industrial dataset provided by Huawei) and eight LLMs, focusing on translation accuracy and safety.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06926
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Integrating Rules and Semantics for LLM-Based C-to-Rust Translation
Luo, Feng
Ji, Kexing
Gao, Cuiyun
Gao, Shuzheng
Feng, Jia
Liu, Kui
Xia, Xin
Lyu, Michael R.
Software Engineering
Automated translation of legacy C code into Rust aims to ensure memory safety while reducing the burden of manual migration. Early approaches in code translation rely on static rule-based methods, but they suffer from limited coverage due to dependence on predefined rule patterns. Recent works regard the task as a sequence-to-sequence problem by leveraging large language models (LLMs). Although these LLM-based methods are capable of reducing unsafe code blocks, the translated code often exhibits issues in following Rust rules and maintaining semantic consistency. On one hand, existing methods adopt a direct prompting strategy to translate the C code, which struggles to accommodate the syntactic rules between C and Rust. On the other hand, this strategy makes it difficult for LLMs to accurately capture the semantics of complex code. To address these challenges, we propose IRENE, an LLM-based framework that Integrates RulEs aNd sEmantics to enhance translation. IRENE consists of three modules: 1) a rule-augmented retrieval module that selects relevant translation examples based on rules generated from a static analyzer developed by us, thereby improving the handling of Rust rules; 2) a structured summarization module that produces a structured summary for guiding LLMs to enhance the semantic understanding of C code; 3) an error-driven translation module that leverages compiler diagnostics to iteratively refine translations. We evaluate IRENE on two datasets (xCodeEval, a public dataset, and HW-Bench, an industrial dataset provided by Huawei) and eight LLMs, focusing on translation accuracy and safety.
title Integrating Rules and Semantics for LLM-Based C-to-Rust Translation
topic Software Engineering
url https://arxiv.org/abs/2508.06926