CODEMENV: Benchmarking Large Language Models on Code Migration

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheng, Keyuan, Shen, Xudong, Yang, Yihao, Wang, Tengyue, Cao, Yang, Ali, Muhammad Asif, Wang, Hanbin, Hu, Lijie, Wang, Di
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910979806199808
author Cheng, Keyuan
Shen, Xudong
Yang, Yihao
Wang, Tengyue
Cao, Yang
Ali, Muhammad Asif
Wang, Hanbin
Hu, Lijie
Wang, Di
author_facet Cheng, Keyuan
Shen, Xudong
Yang, Yihao
Wang, Tengyue
Cao, Yang
Ali, Muhammad Asif
Wang, Hanbin
Hu, Lijie
Wang, Di
contents Large language models (LLMs) have shown remarkable capabilities across various software engineering tasks; however, their effectiveness in code migration, adapting code to run in different environments, remains insufficiently studied. In this work, we introduce CODEMENV: Code Migration Across Environment, a new benchmark specifically designed to assess LLMs' abilities in code migration scenarios. CODEMENV consists of 922 examples spanning 19 Python and Java packages, and covers three core tasks: (1) identifying functions incompatible with specific versions, (2) detecting changes in function definitions, and (3) adapting code to target environments. Experimental evaluation with seven LLMs on CODEMENV yields an average pass@1 rate of 26.50%, with GPT-4O achieving the highest score at 43.84%. Key findings include: (i) LLMs tend to be more proficient with newer function versions, which aids in migrating legacy code, and (ii) LLMs sometimes exhibit logical inconsistencies by identifying function changes irrelevant to the intended migration environment. The datasets are available at https://github.com/xdshen-ai/Benchmark-of-Code-Migration.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00894
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CODEMENV: Benchmarking Large Language Models on Code Migration
Cheng, Keyuan
Shen, Xudong
Yang, Yihao
Wang, Tengyue
Cao, Yang
Ali, Muhammad Asif
Wang, Hanbin
Hu, Lijie
Wang, Di
Software Engineering
Artificial Intelligence
Computation and Language
Machine Learning
Large language models (LLMs) have shown remarkable capabilities across various software engineering tasks; however, their effectiveness in code migration, adapting code to run in different environments, remains insufficiently studied. In this work, we introduce CODEMENV: Code Migration Across Environment, a new benchmark specifically designed to assess LLMs' abilities in code migration scenarios. CODEMENV consists of 922 examples spanning 19 Python and Java packages, and covers three core tasks: (1) identifying functions incompatible with specific versions, (2) detecting changes in function definitions, and (3) adapting code to target environments. Experimental evaluation with seven LLMs on CODEMENV yields an average pass@1 rate of 26.50%, with GPT-4O achieving the highest score at 43.84%. Key findings include: (i) LLMs tend to be more proficient with newer function versions, which aids in migrating legacy code, and (ii) LLMs sometimes exhibit logical inconsistencies by identifying function changes irrelevant to the intended migration environment. The datasets are available at https://github.com/xdshen-ai/Benchmark-of-Code-Migration.
title CODEMENV: Benchmarking Large Language Models on Code Migration
topic Software Engineering
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.00894