Asm2SrcEval: Evaluating Large Language Models for Assembly-to-Source Code Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hamedi, Parisa, Jelodar, Hamed, Bai, Samita, Meymani, Mohammad, Razavi-Far, Roozbeh, Ghorbani, Ali A.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914174464950272
author Hamedi, Parisa
Jelodar, Hamed
Bai, Samita
Meymani, Mohammad
Razavi-Far, Roozbeh
Ghorbani, Ali A.
author_facet Hamedi, Parisa
Jelodar, Hamed
Bai, Samita
Meymani, Mohammad
Razavi-Far, Roozbeh
Ghorbani, Ali A.
contents Assembly-to-source code translation is a critical task in reverse engineering, cybersecurity, and software maintenance, yet systematic benchmarks for evaluating large language models on this problem remain scarce. In this work, we present the first comprehensive evaluation of five state-of-the-art large language models on assembly-to-source translation. We assess model performance using a diverse set of metrics capturing lexical similarity (BLEU, ROUGE, and METEOR), semantic alignment (BERTScore), fluency (Perplexity), and efficiency (time prediction). Our results reveal clear trade-offs: while certain models excel in text similarity metrics, others demonstrate lower perplexity or faster inference times. We further provide qualitative analyses of typical model successes and failure cases, highlighting challenges such as control flow recovery and identifier reconstruction. Taken together, our benchmark offers actionable insights into the strengths and limitations of current large language models for program translation, establishing a foundation for future research in combining accuracy with efficiency for real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00134
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Asm2SrcEval: Evaluating Large Language Models for Assembly-to-Source Code Translation
Hamedi, Parisa
Jelodar, Hamed
Bai, Samita
Meymani, Mohammad
Razavi-Far, Roozbeh
Ghorbani, Ali A.
Software Engineering
Artificial Intelligence
Assembly-to-source code translation is a critical task in reverse engineering, cybersecurity, and software maintenance, yet systematic benchmarks for evaluating large language models on this problem remain scarce. In this work, we present the first comprehensive evaluation of five state-of-the-art large language models on assembly-to-source translation. We assess model performance using a diverse set of metrics capturing lexical similarity (BLEU, ROUGE, and METEOR), semantic alignment (BERTScore), fluency (Perplexity), and efficiency (time prediction). Our results reveal clear trade-offs: while certain models excel in text similarity metrics, others demonstrate lower perplexity or faster inference times. We further provide qualitative analyses of typical model successes and failure cases, highlighting challenges such as control flow recovery and identifier reconstruction. Taken together, our benchmark offers actionable insights into the strengths and limitations of current large language models for program translation, establishing a foundation for future research in combining accuracy with efficiency for real-world applications.
title Asm2SrcEval: Evaluating Large Language Models for Assembly-to-Source Code Translation
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2512.00134