Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maveli, Nickil, Vergari, Antonio, Cohen, Shay B.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910186760830976
author Maveli, Nickil
Vergari, Antonio
Cohen, Shay B.
author_facet Maveli, Nickil
Vergari, Antonio
Cohen, Shay B.
contents LLMs demonstrate strong performance on code benchmarks, yet consistent reasoning across forward and backward execution remains elusive. We present RoundTripCodeEval (RTCE), a benchmark of four code execution reasoning tasks that evaluates round-trip consistency through execution-free, exact-match assessment of bijection fidelity across four lossless compression algorithms. We evaluate state-of-the-art Code-LLMs under zero-shot prompting, supervised fine-tuning on execution traces, and iterative self-reflection. All approaches yield only modest improvements and none closes the gap, revealing that current LLMs lack the internal coherence required for reliable bidirectional code reasoning. RTCE surfaces findings invisible to existing benchmarks: models frequently pass individual forward and backward tasks yet fail the combined round-trip, exposing mutually inconsistent internal representations; SFT and self-reflection saturate after one revision round, indicating they cannot repair fundamental algorithmic misunderstandings; and failures persist even on simple bijections such as RLE, suggesting that algorithmic complexity is not the sole root cause.\footnote{Code and dataset are available at https://github.com/Nickil21/round-trip-code-compression.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13398
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
Maveli, Nickil
Vergari, Antonio
Cohen, Shay B.
Machine Learning
Artificial Intelligence
Programming Languages
LLMs demonstrate strong performance on code benchmarks, yet consistent reasoning across forward and backward execution remains elusive. We present RoundTripCodeEval (RTCE), a benchmark of four code execution reasoning tasks that evaluates round-trip consistency through execution-free, exact-match assessment of bijection fidelity across four lossless compression algorithms. We evaluate state-of-the-art Code-LLMs under zero-shot prompting, supervised fine-tuning on execution traces, and iterative self-reflection. All approaches yield only modest improvements and none closes the gap, revealing that current LLMs lack the internal coherence required for reliable bidirectional code reasoning. RTCE surfaces findings invisible to existing benchmarks: models frequently pass individual forward and backward tasks yet fail the combined round-trip, exposing mutually inconsistent internal representations; SFT and self-reflection saturate after one revision round, indicating they cannot repair fundamental algorithmic misunderstandings; and failures persist even on simple bijections such as RLE, suggesting that algorithmic complexity is not the sole root cause.\footnote{Code and dataset are available at https://github.com/Nickil21/round-trip-code-compression.
title Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
topic Machine Learning
Artificial Intelligence
Programming Languages
url https://arxiv.org/abs/2601.13398