QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ke, Changxin, Zhang, Rui, Wang, Shuo, Ding, Li, Li, Guangli, Wen, Yuanbo, Zhang, Shuoming, Xu, Ruiyuan, Qin, Jin, Guo, Jiaming, Wang, Chenxi, Li, Ling, Guo, Qi, Chen, Yunji
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909863998652416
author Ke, Changxin
Zhang, Rui
Wang, Shuo
Ding, Li
Li, Guangli
Wen, Yuanbo
Zhang, Shuoming
Xu, Ruiyuan
Qin, Jin
Guo, Jiaming
Wang, Chenxi
Li, Ling
Guo, Qi
Chen, Yunji
author_facet Ke, Changxin
Zhang, Rui
Wang, Shuo
Ding, Li
Li, Guangli
Wen, Yuanbo
Zhang, Shuoming
Xu, Ruiyuan
Qin, Jin
Guo, Jiaming
Wang, Chenxi
Li, Ling
Guo, Qi
Chen, Yunji
contents The rise of GPU-based high-performance computing (HPC) has driven the widespread adoption of parallel programming models such as CUDA. Yet, the inherent complexity of parallel programming creates a demand for the automated sequential-to-parallel approaches. However, data scarcity poses a significant challenge for machine learning-based sequential-to-parallel code translation. Although recent back-translation methods show promise, they still fail to ensure functional equivalence in the translated code. In this paper, we propose \textbf{QiMeng-MuPa}, a novel \textbf{Mu}tual-Supervised Learning framework for Sequential-to-\textbf{Pa}rallel code translation, to address the functional equivalence issue. QiMeng-MuPa consists of two models, a Translator and a Tester. Through an iterative loop consisting of Co-verify and Co-evolve steps, the Translator and the Tester mutually generate data for each other and improve collectively. The Tester generates unit tests to verify and filter functionally equivalent translated code, thereby evolving the Translator, while the Translator generates translated code as augmented input to evolve the Tester. Experimental results demonstrate that QiMeng-MuPa significantly enhances the performance of the base models: when applied to Qwen2.5-Coder, it not only improves Pass@1 by up to 28.91% and boosts Tester performance by 68.90%, but also outperforms the previous state-of-the-art method CodeRosetta by 1.56 and 6.92 in BLEU and CodeBLEU scores, while achieving performance comparable to DeepSeek-R1 and GPT-4.1. Our code is available at https://github.com/kcxain/mupa.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11153
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
Ke, Changxin
Zhang, Rui
Wang, Shuo
Ding, Li
Li, Guangli
Wen, Yuanbo
Zhang, Shuoming
Xu, Ruiyuan
Qin, Jin
Guo, Jiaming
Wang, Chenxi
Li, Ling
Guo, Qi
Chen, Yunji
Software Engineering
Machine Learning
The rise of GPU-based high-performance computing (HPC) has driven the widespread adoption of parallel programming models such as CUDA. Yet, the inherent complexity of parallel programming creates a demand for the automated sequential-to-parallel approaches. However, data scarcity poses a significant challenge for machine learning-based sequential-to-parallel code translation. Although recent back-translation methods show promise, they still fail to ensure functional equivalence in the translated code. In this paper, we propose \textbf{QiMeng-MuPa}, a novel \textbf{Mu}tual-Supervised Learning framework for Sequential-to-\textbf{Pa}rallel code translation, to address the functional equivalence issue. QiMeng-MuPa consists of two models, a Translator and a Tester. Through an iterative loop consisting of Co-verify and Co-evolve steps, the Translator and the Tester mutually generate data for each other and improve collectively. The Tester generates unit tests to verify and filter functionally equivalent translated code, thereby evolving the Translator, while the Translator generates translated code as augmented input to evolve the Tester. Experimental results demonstrate that QiMeng-MuPa significantly enhances the performance of the base models: when applied to Qwen2.5-Coder, it not only improves Pass@1 by up to 28.91% and boosts Tester performance by 68.90%, but also outperforms the previous state-of-the-art method CodeRosetta by 1.56 and 6.92 in BLEU and CodeBLEU scores, while achieving performance comparable to DeepSeek-R1 and GPT-4.1. Our code is available at https://github.com/kcxain/mupa.
title QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2506.11153