Toward Explaining Large Language Models in Software Engineering Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Vitale, Antonio, Nguyen, Khai-Nguyen, Poshyvanyk, Denys, Oliveto, Rocco, Scalabrino, Simone, Mastropaolo, Antonio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
by: Vitale, Antonio, et al.
Published: (2026)
by: Vitale, Antonio, et al.
Published: (2026)
A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm
by: Mastropaolo, Antonio, et al.
Published: (2025)
by: Mastropaolo, Antonio, et al.
Published: (2025)
Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?
by: Vitale, Antonio, et al.
Published: (2025)
by: Vitale, Antonio, et al.
Published: (2025)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
Toward Neurosymbolic Program Comprehension
by: Velasco, Alejandro, et al.
Published: (2025)
by: Velasco, Alejandro, et al.
Published: (2025)
Personalized Code Readability Assessment: Are We There Yet?
by: Vitale, Antonio, et al.
Published: (2025)
by: Vitale, Antonio, et al.
Published: (2025)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2025)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2025)
Toward a Theory of Causation for Interpreting Neural Code Models
by: Palacio, David N., et al.
Published: (2023)
by: Palacio, David N., et al.
Published: (2023)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
by: Palacio, David N., et al.
Published: (2024)
by: Palacio, David N., et al.
Published: (2024)
Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives
by: Khati, Dipin, et al.
Published: (2025)
by: Khati, Dipin, et al.
Published: (2025)
Challenges and Paths Towards AI for Software Engineering
by: Gu, Alex, et al.
Published: (2025)
by: Gu, Alex, et al.
Published: (2025)
"Don't Be Afraid, Just Learn": Insights from Industry Practitioners to Prepare Software Engineers in the Age of Generative AI
by: Otten, Daniel, et al.
Published: (2026)
by: Otten, Daniel, et al.
Published: (2026)
Towards a Classification of Open-Source ML Models and Datasets for Software Engineering
by: González, Alexandra, et al.
Published: (2024)
by: González, Alexandra, et al.
Published: (2024)
On Interpreting the Effectiveness of Unsupervised Software Traceability with Information Theory
by: Palacio, David N., et al.
Published: (2024)
by: Palacio, David N., et al.
Published: (2024)
Foundation Model Engineering: Engineering Foundation Models Just as Engineering Software
by: Ran, Dezhi, et al.
Published: (2024)
by: Ran, Dezhi, et al.
Published: (2024)
Perspective of Software Engineering Researchers on Machine Learning Practices Regarding Research, Review, and Education
by: Mojica-Hanke, Anamaria, et al.
Published: (2024)
by: Mojica-Hanke, Anamaria, et al.
Published: (2024)
Beyond Code Similarity: Benchmarking the Plausibility, Efficiency, and Complexity of LLM-Generated Smart Contracts
by: Salzano, Francesco, et al.
Published: (2025)
by: Salzano, Francesco, et al.
Published: (2025)
Fixing Smart Contract Vulnerabilities: A Comparative Analysis of Literature and Developer's Practices
by: Salzano, Francesco, et al.
Published: (2024)
by: Salzano, Francesco, et al.
Published: (2024)
The Rise and Fall(?) of Software Engineering
by: Mastropaolo, Antonio, et al.
Published: (2024)
by: Mastropaolo, Antonio, et al.
Published: (2024)
How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study
by: Velasco, Alejandro, et al.
Published: (2024)
by: Velasco, Alejandro, et al.
Published: (2024)
Tricky$^2$: Towards a Benchmark for Evaluating Human and LLM Error Interactions
by: Granger, Cole, et al.
Published: (2026)
by: Granger, Cole, et al.
Published: (2026)
Benchmarking Large Language Models with Integer Sequence Generation Tasks
by: O'Malley, Daniel, et al.
Published: (2024)
by: O'Malley, Daniel, et al.
Published: (2024)
Towards Engineering Fair and Equitable Software Systems for Managing Low-Altitude Airspace Authorizations
by: Gohar, Usman, et al.
Published: (2024)
by: Gohar, Usman, et al.
Published: (2024)
Explaining Software Vulnerabilities with Large Language Models
by: Johnson, Oshando, et al.
Published: (2025)
by: Johnson, Oshando, et al.
Published: (2025)
HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale
by: Phan, Huy Nhat, et al.
Published: (2024)
by: Phan, Huy Nhat, et al.
Published: (2024)
Order Matters! An Empirical Study on Large Language Models' Input Order Bias in Software Fault Localization
by: Rafi, Md Nakhla, et al.
Published: (2024)
by: Rafi, Md Nakhla, et al.
Published: (2024)
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
by: Zhou, Yongxi, et al.
Published: (2026)
by: Zhou, Yongxi, et al.
Published: (2026)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
by: Li, Yuangang, et al.
Published: (2026)
by: Li, Yuangang, et al.
Published: (2026)
On the Replicability and Reproducibility of Deep Learning in Software Engineering
by: Liu, Chao, et al.
Published: (2020)
by: Liu, Chao, et al.
Published: (2020)
LLMs in Coding and their Impact on the Commercial Software Engineering Landscape
by: Belozerov, Vladislav, et al.
Published: (2025)
by: Belozerov, Vladislav, et al.
Published: (2025)
EGSS: Entropy-guided Stepwise Scaling for Reliable Software Engineering
by: Mao, Chenhui, et al.
Published: (2026)
by: Mao, Chenhui, et al.
Published: (2026)
A Large-Scale Study of Model Integration in ML-Enabled Software Systems
by: Sens, Yorick, et al.
Published: (2024)
by: Sens, Yorick, et al.
Published: (2024)
How do Copilot Suggestions Impact Developers' Frustration and Productivity?
by: Guglielmi, Emanuela, et al.
Published: (2025)
by: Guglielmi, Emanuela, et al.
Published: (2025)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
by: Ding, Yifeng, et al.
Published: (2026)
by: Ding, Yifeng, et al.
Published: (2026)
Rethinking Software Empirical Studies with Structural Causal Models
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
by: Khati, Dipin, et al.
Published: (2026)
by: Khati, Dipin, et al.
Published: (2026)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
by: Arora, Avi, et al.
Published: (2025)
by: Arora, Avi, et al.
Published: (2025)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
by: Yuan, Danlong, et al.
Published: (2026)
by: Yuan, Danlong, et al.
Published: (2026)
Analysing the Behaviour of Tree-Based Neural Networks in Regression Tasks
by: Samoaa, Peter, et al.
Published: (2024)
by: Samoaa, Peter, et al.
Published: (2024)
DRAGON: Robust Classification for Very Large Collections of Software Repositories
by: Balla, Stefano, et al.
Published: (2026)
by: Balla, Stefano, et al.
Published: (2026)
Similar Items
-
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
by: Vitale, Antonio, et al.
Published: (2026) -
A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm
by: Mastropaolo, Antonio, et al.
Published: (2025) -
Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?
by: Vitale, Antonio, et al.
Published: (2025) -
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026) -
Toward Neurosymbolic Program Comprehension
by: Velasco, Alejandro, et al.
Published: (2025)