Large Language Models for Software Engineering: A Reproducibility Crisis
Fuente:
arXiv
Saved in:
| Main Authors: | Siddiq, Mohammed Latif, Islam-Gomes, Arvin, Sekerak, Natalie, Santos, Joanna C. S. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assessing the Software Security Comprehension of Large Language Models
by: Siddiq, Mohammed Latif, et al.
Published: (2025)
by: Siddiq, Mohammed Latif, et al.
Published: (2025)
FRANC: A Lightweight Framework for High-Quality Code Generation
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
An Empirical Study on Remote Code Execution in Machine Learning Model Hosting Ecosystems
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
by: Siddiq, Mohammed Latif, et al.
Published: (2024)
by: Siddiq, Mohammed Latif, et al.
Published: (2024)
Using Large Language Models to Generate JUnit Tests: An Empirical Study
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
More Rigorous Software Engineering Would Improve Reproducibility in Machine Learning Research
by: Wolter, Moritz, et al.
Published: (2025)
by: Wolter, Moritz, et al.
Published: (2025)
On the Replicability and Reproducibility of Deep Learning in Software Engineering
by: Liu, Chao, et al.
Published: (2020)
by: Liu, Chao, et al.
Published: (2020)
SALLM: Security Assessment of Generated Code
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
by: Siddiq, Mohammed Latif, et al.
Published: (2023)
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
by: Tao, Hongyuan, et al.
Published: (2025)
by: Tao, Hongyuan, et al.
Published: (2025)
IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks
by: Mateega, Spencer, et al.
Published: (2026)
by: Mateega, Spencer, et al.
Published: (2026)
Toward Explaining Large Language Models in Software Engineering Tasks
by: Vitale, Antonio, et al.
Published: (2025)
by: Vitale, Antonio, et al.
Published: (2025)
Investigating Reproducibility in Deep Learning-Based Software Fault Prediction
by: Mukhtar, Adil, et al.
Published: (2024)
by: Mukhtar, Adil, et al.
Published: (2024)
Improving the Reproducibility of Deep Learning Software: An Initial Investigation through a Case Study Analysis
by: Ravi, Nikita, et al.
Published: (2025)
by: Ravi, Nikita, et al.
Published: (2025)
Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility
by: Jin, Bihui, et al.
Published: (2026)
by: Jin, Bihui, et al.
Published: (2026)
Combating Toxic Language: A Review of LLM-Based Strategies for Software Engineering
by: Zhuo, Hao, et al.
Published: (2025)
by: Zhuo, Hao, et al.
Published: (2025)
Unlearning Trojans in Large Language Models: A Comparison Between Natural Language and Source Code
by: Kazemi, Mahdi, et al.
Published: (2024)
by: Kazemi, Mahdi, et al.
Published: (2024)
Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
by: Siddiq, Mohammed Latif, et al.
Published: (2026)
Effort and Size Estimation in Software Projects with Large Language Model-based Intelligent Interfaces
by: Coelho Jr, Claudionor N., et al.
Published: (2024)
by: Coelho Jr, Claudionor N., et al.
Published: (2024)
A Large-Scale Exploit Instrumentation Study of AI/ML Supply Chain Attacks in Hugging Face Models
by: Casey, Beatrice, et al.
Published: (2024)
by: Casey, Beatrice, et al.
Published: (2024)
Cataloguing Hugging Face Models to Software Engineering Activities: Automation and Findings
by: González, Alexandra, et al.
Published: (2025)
by: González, Alexandra, et al.
Published: (2025)
SWE-Arena: An Interactive Platform for Evaluating Foundation Models in Software Engineering
by: Zhao, Zhimin
Published: (2025)
by: Zhao, Zhimin
Published: (2025)
Foundation Model Engineering: Engineering Foundation Models Just as Engineering Software
by: Ran, Dezhi, et al.
Published: (2024)
by: Ran, Dezhi, et al.
Published: (2024)
Applying the Chinese Wall Reverse Engineering Technique to Large Language Model Code Editing
by: Hanmongkolchai, Manatsawin
Published: (2025)
by: Hanmongkolchai, Manatsawin
Published: (2025)
Trojans in Large Language Models of Code: A Critical Review through a Trigger-Based Taxonomy
by: Hussain, Aftab, et al.
Published: (2024)
by: Hussain, Aftab, et al.
Published: (2024)
Agint: Agentic Graph Compilation for Software Engineering Agents
by: Chivukula, Abhi, et al.
Published: (2025)
by: Chivukula, Abhi, et al.
Published: (2025)
Breaking the Silence: the Threats of Using LLMs in Software Engineering
by: Sallou, June, et al.
Published: (2023)
by: Sallou, June, et al.
Published: (2023)
A Systematic Literature Review on the Use of Machine Learning in Software Engineering
by: Fred, Nyaga, et al.
Published: (2024)
by: Fred, Nyaga, et al.
Published: (2024)
Software Engineering Principles for Fairer Systems: Experiments with GroupCART
by: Peng, Kewen, et al.
Published: (2025)
by: Peng, Kewen, et al.
Published: (2025)
Assessing the Use of AutoML for Data-Driven Software Engineering
by: Calefato, Fabio, et al.
Published: (2023)
by: Calefato, Fabio, et al.
Published: (2023)
Should Code Models Learn Pedagogically? A Preliminary Evaluation of Curriculum Learning for Real-World Software Engineering Tasks
by: Khant, Kyi Shin, et al.
Published: (2025)
by: Khant, Kyi Shin, et al.
Published: (2025)
On Trojan Signatures in Large Language Models of Code
by: Hussain, Aftab, et al.
Published: (2024)
by: Hussain, Aftab, et al.
Published: (2024)
Applying Large Language Models to Issue Classification: Revisiting with Extended Data and New Models
by: Aracena, Gabriel, et al.
Published: (2025)
by: Aracena, Gabriel, et al.
Published: (2025)
Engineering Resource-constrained Software Systems with DNN Components: a Concept-based Pruning Approach
by: Formica, Federico, et al.
Published: (2026)
by: Formica, Federico, et al.
Published: (2026)
An ML-based Approach to Predicting Software Change Dependencies: Insights from an Empirical Study on OpenStack
by: Arabat, Ali, et al.
Published: (2025)
by: Arabat, Ali, et al.
Published: (2025)
SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow
by: Bula, Timothy, et al.
Published: (2025)
by: Bula, Timothy, et al.
Published: (2025)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
by: Aggarwal, Pranjal, et al.
Published: (2025)
by: Aggarwal, Pranjal, et al.
Published: (2025)
A Configuration-First Framework for Reproducible, Low-Code Localization
by: Strnad, Tim, et al.
Published: (2025)
by: Strnad, Tim, et al.
Published: (2025)
Program Structure-aware Language Models: Targeted Software Testing beyond Textual Semantics
by: Tran, Khang, et al.
Published: (2026)
by: Tran, Khang, et al.
Published: (2026)
Perspective of Software Engineering Researchers on Machine Learning Practices Regarding Research, Review, and Education
by: Mojica-Hanke, Anamaria, et al.
Published: (2024)
by: Mojica-Hanke, Anamaria, et al.
Published: (2024)
R-LAM: Reproducibility-Constrained Large Action Models for Scientific Workflow Automation
by: Sureshkumar, Suriya
Published: (2026)
by: Sureshkumar, Suriya
Published: (2026)
Similar Items
-
Assessing the Software Security Comprehension of Large Language Models
by: Siddiq, Mohammed Latif, et al.
Published: (2025) -
FRANC: A Lightweight Framework for High-Quality Code Generation
by: Siddiq, Mohammed Latif, et al.
Published: (2023) -
An Empirical Study on Remote Code Execution in Machine Learning Model Hosting Ecosystems
by: Siddiq, Mohammed Latif, et al.
Published: (2026) -
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
by: Siddiq, Mohammed Latif, et al.
Published: (2024) -
Using Large Language Models to Generate JUnit Tests: An Empirical Study
by: Siddiq, Mohammed Latif, et al.
Published: (2023)