Assessing Code Understanding in LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Laneve, Cosimo, Spanò, Alvise, Ressi, Dalila, Rossi, Sabina, Bugliesi, Michele
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913769141043200
author Laneve, Cosimo
Spanò, Alvise
Ressi, Dalila
Rossi, Sabina
Bugliesi, Michele
author_facet Laneve, Cosimo
Spanò, Alvise
Ressi, Dalila
Rossi, Sabina
Bugliesi, Michele
contents We present an empirical evaluation of Large Language Models in code understanding associated with non-trivial, semantic-preserving program transformations such as copy propagation or constant folding. Our findings show that LLMs fail to judge semantic equivalence in approximately 41\% of cases when no context is provided and in 29\% when given a simple generic context. To improve accuracy, we advocate integrating LLMs with code-optimization tools to enhance training and facilitate more robust program understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00065
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing Code Understanding in LLMs
Laneve, Cosimo
Spanò, Alvise
Ressi, Dalila
Rossi, Sabina
Bugliesi, Michele
Software Engineering
Artificial Intelligence
Programming Languages
We present an empirical evaluation of Large Language Models in code understanding associated with non-trivial, semantic-preserving program transformations such as copy propagation or constant folding. Our findings show that LLMs fail to judge semantic equivalence in approximately 41\% of cases when no context is provided and in 29\% when given a simple generic context. To improve accuracy, we advocate integrating LLMs with code-optimization tools to enhance training and facilitate more robust program understanding.
title Assessing Code Understanding in LLMs
topic Software Engineering
Artificial Intelligence
Programming Languages
url https://arxiv.org/abs/2504.00065