How Robustly do LLMs Understand Execution Semantics?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Spiess, Claudio, Devanbu, Prem, Barr, Earl T. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
On LLMs' Internal Representation of Code Correctness
par: Ribeiro, Francisco, et autres
Publié: (2025)
par: Ribeiro, Francisco, et autres
Publié: (2025)
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
par: Al-Kaswan, Ali, et autres
Publié: (2026)
par: Al-Kaswan, Ali, et autres
Publié: (2026)
Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)
par: Ahmed, Toufique, et autres
Publié: (2023)
par: Ahmed, Toufique, et autres
Publié: (2023)
Localized Calibrated Uncertainty in Code Language Models
par: Gros, David, et autres
Publié: (2025)
par: Gros, David, et autres
Publié: (2025)
Does In-IDE Calibration of Large Language Models work at Scale?
par: Koohestani, Roham, et autres
Publié: (2025)
par: Koohestani, Roham, et autres
Publié: (2025)
Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time
par: Popescu, Razvan Mihai, et autres
Publié: (2026)
par: Popescu, Razvan Mihai, et autres
Publié: (2026)
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
par: Gu, Alex, et autres
Publié: (2024)
par: Gu, Alex, et autres
Publié: (2024)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
par: Jiang, Shan, et autres
Publié: (2026)
par: Jiang, Shan, et autres
Publié: (2026)
Operational Robustness of LLMs on Code Generation
par: Paul, Debalina Ghosh, et autres
Publié: (2026)
par: Paul, Debalina Ghosh, et autres
Publié: (2026)
Calibration and Correctness of Language Models for Code
par: Spiess, Claudio, et autres
Publié: (2024)
par: Spiess, Claudio, et autres
Publié: (2024)
Towards Understanding What Code Language Models Learned
par: Ahmed, Toufique, et autres
Publié: (2023)
par: Ahmed, Toufique, et autres
Publié: (2023)
CAPE: Capability Achievement via Policy Execution
par: Ball, David
Publié: (2025)
par: Ball, David
Publié: (2025)
How Robust are LLM-Generated Library Imports? An Empirical Study using Stack Overflow
par: Latendresse, Jasmine, et autres
Publié: (2025)
par: Latendresse, Jasmine, et autres
Publié: (2025)
Automatic Generation of Executable BPMN Models from Medical Guidelines
par: Sekar, Praveen Kumar Menaka, et autres
Publié: (2026)
par: Sekar, Praveen Kumar Menaka, et autres
Publié: (2026)
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
par: Liu, Hugh Xuechen, et autres
Publié: (2026)
par: Liu, Hugh Xuechen, et autres
Publié: (2026)
Feedback Over Form: Why Execution Feedback Matters More Than Pipeline Topology in 1-3B Code Generation
par: McAndrews, Charles Junichi
Publié: (2026)
par: McAndrews, Charles Junichi
Publié: (2026)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
par: Bouzenia, Islem, et autres
Publié: (2024)
par: Bouzenia, Islem, et autres
Publié: (2024)
Understanding LLM-Driven Test Oracle Generation
par: Bodicoat, Adam, et autres
Publié: (2026)
par: Bodicoat, Adam, et autres
Publié: (2026)
SPELL: Synthesis of Programmatic Edits using LLMs
par: Ramos, Daniel, et autres
Publié: (2026)
par: Ramos, Daniel, et autres
Publié: (2026)
Evaluating the Use of LLMs for Documentation to Code Traceability
par: Alor, Ebube, et autres
Publié: (2025)
par: Alor, Ebube, et autres
Publié: (2025)
Methodological Framework for Quantifying Semantic Test Coverage in RAG Systems
par: Broestl, Noah, et autres
Publié: (2025)
par: Broestl, Noah, et autres
Publié: (2025)
DRAFT-ing Architectural Design Decisions using LLMs
par: Dhar, Rudra, et autres
Publié: (2025)
par: Dhar, Rudra, et autres
Publié: (2025)
LLMs in Coding and their Impact on the Commercial Software Engineering Landscape
par: Belozerov, Vladislav, et autres
Publié: (2025)
par: Belozerov, Vladislav, et autres
Publié: (2025)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
par: Bodla, Krishna Vamshi, et autres
Publié: (2025)
par: Bodla, Krishna Vamshi, et autres
Publié: (2025)
The Struggles of LLMs in Cross-lingual Code Clone Detection
par: Moumoula, Micheline Bénédicte, et autres
Publié: (2024)
par: Moumoula, Micheline Bénédicte, et autres
Publié: (2024)
Machine Learning Robustness: A Primer
par: Braiek, Houssem Ben, et autres
Publié: (2024)
par: Braiek, Houssem Ben, et autres
Publié: (2024)
On Wasted Contributions: Understanding the Dynamics of Contributor-Abandoned Pull Requests
par: Khatoonabadi, SayedHassan, et autres
Publié: (2021)
par: Khatoonabadi, SayedHassan, et autres
Publié: (2021)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
par: Patel, Harsh, et autres
Publié: (2024)
par: Patel, Harsh, et autres
Publié: (2024)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
par: Galimzyanov, Timur, et autres
Publié: (2024)
par: Galimzyanov, Timur, et autres
Publié: (2024)
A Regression Framework for Understanding Prompt Component Impact on LLM Performance
par: Lauziere, Andrew, et autres
Publié: (2026)
par: Lauziere, Andrew, et autres
Publié: (2026)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
par: Vulićević, Jelena Ilić
Publié: (2026)
par: Vulićević, Jelena Ilić
Publié: (2026)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
par: Zhou, Zenghui, et autres
Publié: (2026)
par: Zhou, Zenghui, et autres
Publié: (2026)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
par: Thillen, Alex, et autres
Publié: (2026)
par: Thillen, Alex, et autres
Publié: (2026)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2025)
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2025)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
par: Xu, Weiwei, et autres
Publié: (2024)
par: Xu, Weiwei, et autres
Publié: (2024)
Free and Customizable Code Documentation with LLMs: A Fine-Tuning Approach
par: Chakrabarty, Sayak, et autres
Publié: (2024)
par: Chakrabarty, Sayak, et autres
Publié: (2024)
Can LLMs Generate Architectural Design Decisions? -An Exploratory Empirical study
par: Dhar, Rudra, et autres
Publié: (2024)
par: Dhar, Rudra, et autres
Publié: (2024)
Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
par: Jahangirova, Gunel, et autres
Publié: (2024)
par: Jahangirova, Gunel, et autres
Publié: (2024)
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
par: Jiralerspong, Thomas, et autres
Publié: (2026)
par: Jiralerspong, Thomas, et autres
Publié: (2026)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
par: Palacio, David N., et autres
Publié: (2024)
par: Palacio, David N., et autres
Publié: (2024)
Documents similaires
-
On LLMs' Internal Representation of Code Correctness
par: Ribeiro, Francisco, et autres
Publié: (2025) -
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
par: Al-Kaswan, Ali, et autres
Publié: (2026) -
Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)
par: Ahmed, Toufique, et autres
Publié: (2023) -
Localized Calibrated Uncertainty in Code Language Models
par: Gros, David, et autres
Publié: (2025) -
Does In-IDE Calibration of Large Language Models work at Scale?
par: Koohestani, Roham, et autres
Publié: (2025)