Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
Fuente:
arXiv
Saved in:
| Main Authors: | Maveli, Nickil, Vergari, Antonio, Cohen, Shay B. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What can Large Language Models Capture about Code Functional Equivalence?
by: Maveli, Nickil, et al.
Published: (2024)
by: Maveli, Nickil, et al.
Published: (2024)
What I cannot execute, I do not understand: Training and Evaluating LLMs on Program Execution Traces
by: Armengol-Estapé, Jordi, et al.
Published: (2025)
by: Armengol-Estapé, Jordi, et al.
Published: (2025)
LILO: Learning Interpretable Libraries by Compressing and Documenting Code
by: Grand, Gabriel, et al.
Published: (2023)
by: Grand, Gabriel, et al.
Published: (2023)
Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions
by: Cassano, Federico, et al.
Published: (2023)
by: Cassano, Federico, et al.
Published: (2023)
Towards LLM-based optimization compilers. Can LLMs learn how to apply a single peephole optimization? Reasoning is all LLMs need!
by: Fang, Xiangxin, et al.
Published: (2024)
by: Fang, Xiangxin, et al.
Published: (2024)
Logically Consistent Language Models via Neuro-Symbolic Integration
by: Calanzone, Diego, et al.
Published: (2024)
by: Calanzone, Diego, et al.
Published: (2024)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
Chain of Execution Supervision Promotes General Reasoning in Large Language Models
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
EnCompass: Enhancing Agent Programming with Search Over Program Execution Paths
by: Li, Zhening, et al.
Published: (2025)
by: Li, Zhening, et al.
Published: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
BaxBench: Can LLMs Generate Correct and Secure Backends?
by: Vero, Mark, et al.
Published: (2025)
by: Vero, Mark, et al.
Published: (2025)
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
by: Hooda, Ashish, et al.
Published: (2024)
by: Hooda, Ashish, et al.
Published: (2024)
Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information
by: Miceli-Barone, Antonio Valerio, et al.
Published: (2026)
by: Miceli-Barone, Antonio Valerio, et al.
Published: (2026)
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
$\textbf{PLUM}$: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases
by: Zhang, Dylan, et al.
Published: (2024)
by: Zhang, Dylan, et al.
Published: (2024)
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
by: Dai, Hankun, et al.
Published: (2025)
by: Dai, Hankun, et al.
Published: (2025)
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
by: Zhao, Zheng, et al.
Published: (2025)
by: Zhao, Zheng, et al.
Published: (2025)
Evaluating LLMs for Hardware Design and Test
by: Blocklove, Jason, et al.
Published: (2024)
by: Blocklove, Jason, et al.
Published: (2024)
Curriculum Learning for Small Code Language Models
by: Naïr, Marwa, et al.
Published: (2024)
by: Naïr, Marwa, et al.
Published: (2024)
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
From Reasoning to Code: GRPO Optimization for Underrepresented Languages
by: Pennino, Federico, et al.
Published: (2025)
by: Pennino, Federico, et al.
Published: (2025)
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
by: Cohen-Inger, Nurit, et al.
Published: (2025)
by: Cohen-Inger, Nurit, et al.
Published: (2025)
Code Simulation Challenges for Large Language Models
by: La Malfa, Emanuele, et al.
Published: (2024)
by: La Malfa, Emanuele, et al.
Published: (2024)
Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis
by: Gajjar, Jugal
Published: (2026)
by: Gajjar, Jugal
Published: (2026)
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
by: Ravi, Nikil, et al.
Published: (2026)
by: Ravi, Nikil, et al.
Published: (2026)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
Bridging the Knowledge Void: Inference-time Acquisition of Unfamiliar Programming Languages for Coding Tasks
by: Shen, Chen, et al.
Published: (2026)
by: Shen, Chen, et al.
Published: (2026)
An Evaluation Benchmark for Autoformalization in Lean4
by: Gulati, Aryan, et al.
Published: (2024)
by: Gulati, Aryan, et al.
Published: (2024)
On Faster Marginalization with Squared Circuits via Orthonormalization
by: Loconte, Lorenzo, et al.
Published: (2024)
by: Loconte, Lorenzo, et al.
Published: (2024)
MonoCoder: Domain-Specific Code Language Model for HPC Codes and Tasks
by: Kadosh, Tal, et al.
Published: (2023)
by: Kadosh, Tal, et al.
Published: (2023)
Operon: Incremental Construction of Ragged Data via Named Dimensions
by: Moon, Sungbin, et al.
Published: (2025)
by: Moon, Sungbin, et al.
Published: (2025)
Program Semantic Inequivalence Game with Large Language Models
by: Miceli-Barone, Antonio Valerio, et al.
Published: (2025)
by: Miceli-Barone, Antonio Valerio, et al.
Published: (2025)
ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning
by: Yang, Chengcao
Published: (2026)
by: Yang, Chengcao
Published: (2026)
From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
by: Gong, Junfeng, et al.
Published: (2025)
by: Gong, Junfeng, et al.
Published: (2025)
Large Language Models for Code Summarization
by: Szalontai, Balázs, et al.
Published: (2024)
by: Szalontai, Balázs, et al.
Published: (2024)
Is Programming by Example solved by LLMs?
by: Li, Wen-Ding, et al.
Published: (2024)
by: Li, Wen-Ding, et al.
Published: (2024)
Spectral Editing of Activations for Large Language Model Alignment
by: Qiu, Yifu, et al.
Published: (2024)
by: Qiu, Yifu, et al.
Published: (2024)
Assessing Code Understanding in LLMs
by: Laneve, Cosimo, et al.
Published: (2025)
by: Laneve, Cosimo, et al.
Published: (2025)
Similar Items
-
What can Large Language Models Capture about Code Functional Equivalence?
by: Maveli, Nickil, et al.
Published: (2024) -
What I cannot execute, I do not understand: Training and Evaluating LLMs on Program Execution Traces
by: Armengol-Estapé, Jordi, et al.
Published: (2025) -
LILO: Learning Interpretable Libraries by Compressing and Documenting Code
by: Grand, Gabriel, et al.
Published: (2023) -
Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions
by: Cassano, Federico, et al.
Published: (2023) -
Towards LLM-based optimization compilers. Can LLMs learn how to apply a single peephole optimization? Reasoning is all LLMs need!
by: Fang, Xiangxin, et al.
Published: (2024)