Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Rodriguez-Cardenas, Daniel, Li, Xiaochang, Macedo, Marcos, Mastropaolo, Antonio, Khati, Dipin, Tian, Yuan, Shao, Huajie, Poshyvanyk, Denys |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Tricky$^2$: Towards a Benchmark for Evaluating Human and LLM Error Interactions
par: Granger, Cole, et autres
Publié: (2026)
par: Granger, Cole, et autres
Publié: (2026)
A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm
par: Mastropaolo, Antonio, et autres
Publié: (2025)
par: Mastropaolo, Antonio, et autres
Publié: (2025)
Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives
par: Khati, Dipin, et autres
Publié: (2025)
par: Khati, Dipin, et autres
Publié: (2025)
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
par: Khati, Dipin, et autres
Publié: (2026)
par: Khati, Dipin, et autres
Publié: (2026)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
par: Palacio, David N., et autres
Publié: (2024)
par: Palacio, David N., et autres
Publié: (2024)
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
par: Khati, Dipin, et autres
Publié: (2025)
par: Khati, Dipin, et autres
Publié: (2025)
A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code
par: Velasco, Alejandro, et autres
Publié: (2025)
par: Velasco, Alejandro, et autres
Publié: (2025)
Toward Neurosymbolic Program Comprehension
par: Velasco, Alejandro, et autres
Publié: (2025)
par: Velasco, Alejandro, et autres
Publié: (2025)
Toward Explaining Large Language Models in Software Engineering Tasks
par: Vitale, Antonio, et autres
Publié: (2025)
par: Vitale, Antonio, et autres
Publié: (2025)
Towards Enabling An Artificial Self-Construction Software Life-cycle via Autopoietic Architectures
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026)
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026)
Rethinking Software Empirical Studies with Structural Causal Models
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026)
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026)
Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution Detection
par: Yan, Yanfu, et autres
Publié: (2025)
par: Yan, Yanfu, et autres
Publié: (2025)
Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We?
par: O'Brien, Conor, et autres
Publié: (2024)
par: O'Brien, Conor, et autres
Publié: (2024)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2025)
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2025)
On Interpreting the Effectiveness of Unsupervised Software Traceability with Information Theory
par: Palacio, David N., et autres
Publié: (2024)
par: Palacio, David N., et autres
Publié: (2024)
Bridging the Quantum Divide: Aligning Academic and Industry Goals in Software Engineering
par: Zappin, Jake, et autres
Publié: (2025)
par: Zappin, Jake, et autres
Publié: (2025)
On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
par: Crupi, Giuseppe, et autres
Publié: (2025)
par: Crupi, Giuseppe, et autres
Publié: (2025)
Which Syntactic Capabilities Are Statistically Learned by Masked Language Models for Code?
par: Velasco, Alejandro, et autres
Publié: (2024)
par: Velasco, Alejandro, et autres
Publié: (2024)
BOMs Away! Inside the Minds of Stakeholders: A Comprehensive Study of Bills of Materials for Software Systems
par: Stalnaker, Trevor, et autres
Publié: (2023)
par: Stalnaker, Trevor, et autres
Publié: (2023)
Challenges and Practices in Quantum Software Testing and Debugging: Insights from Practitioners
par: Zappin, Jake, et autres
Publié: (2025)
par: Zappin, Jake, et autres
Publié: (2025)
Prompting in Practice: Investigating Software Practitioners' Use of Generative AI Tools
par: Otten, Daniel, et autres
Publié: (2025)
par: Otten, Daniel, et autres
Publié: (2025)
How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study
par: Velasco, Alejandro, et autres
Publié: (2024)
par: Velasco, Alejandro, et autres
Publié: (2024)
"Don't Be Afraid, Just Learn": Insights from Industry Practitioners to Prepare Software Engineers in the Age of Generative AI
par: Otten, Daniel, et autres
Publié: (2026)
par: Otten, Daniel, et autres
Publié: (2026)
Smart but Costly? Benchmarking LLMs on Functional Accuracy and Energy Efficiency
par: Mehditabar, Mohammadjavad, et autres
Publié: (2025)
par: Mehditabar, Mohammadjavad, et autres
Publié: (2025)
Testing Practices, Challenges, and Developer Perspectives in Open-Source IoT Platforms
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2025)
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2025)
Perspective of Software Engineering Researchers on Machine Learning Practices Regarding Research, Review, and Education
par: Mojica-Hanke, Anamaria, et autres
Publié: (2024)
par: Mojica-Hanke, Anamaria, et autres
Publié: (2024)
"The Law Doesn't Work Like a Computer": Exploring Software Licensing Issues Faced by Legal Practitioners
par: Wintersgill, Nathan, et autres
Publié: (2024)
par: Wintersgill, Nathan, et autres
Publié: (2024)
When Quantum Meets Classical: Characterizing Hybrid Quantum-Classical Issues Discussed in Developer Forums
par: Zappin, Jake, et autres
Publié: (2024)
par: Zappin, Jake, et autres
Publié: (2024)
On the Generalizability of Transformer Models to Code Completions of Different Lengths
par: Cooper, Nathan, et autres
Publié: (2025)
par: Cooper, Nathan, et autres
Publié: (2025)
The Rise and Fall(?) of Software Engineering
par: Mastropaolo, Antonio, et autres
Publié: (2024)
par: Mastropaolo, Antonio, et autres
Publié: (2024)
Developers' Perspectives on Software Licensing: Current Practices, Challenges, and Tools
par: Wintersgill, Nathan, et autres
Publié: (2025)
par: Wintersgill, Nathan, et autres
Publié: (2025)
An Empirical Study on the Effects of System Prompts in Instruction-Tuned Models for Code Generation
par: Cheng, Zaiyu, et autres
Publié: (2026)
par: Cheng, Zaiyu, et autres
Publié: (2026)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
par: Fang, Sen, et autres
Publié: (2025)
par: Fang, Sen, et autres
Publié: (2025)
Investigating the Use of LLMs for Evidence Briefings Generation in Software Engineering
par: Marcelino, Mauro, et autres
Publié: (2025)
par: Marcelino, Mauro, et autres
Publié: (2025)
Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development
par: Stalnaker, Trevor, et autres
Publié: (2024)
par: Stalnaker, Trevor, et autres
Publié: (2024)
Teaching Survey Research in Software Engineering
par: Kalinowski, Marcos, et autres
Publié: (2024)
par: Kalinowski, Marcos, et autres
Publié: (2024)
Is Quantization a Deal-breaker? Empirical Insights from Large Code Models
par: Afrin, Saima, et autres
Publié: (2025)
par: Afrin, Saima, et autres
Publié: (2025)
"False negative -- that one is going to kill you": Understanding Industry Perspectives of Static Analysis based Security Testing
par: Ami, Amit Seal, et autres
Publié: (2023)
par: Ami, Amit Seal, et autres
Publié: (2023)
SENAI: Towards Software Engineering Native Generative Artificial Intelligence
par: Saad, Mootez, et autres
Publié: (2025)
par: Saad, Mootez, et autres
Publié: (2025)
An Infrastructure Software Perspective Toward Computation Offloading between Executable Specifications and Foundation Models
par: Ran, Dezhi, et autres
Publié: (2025)
par: Ran, Dezhi, et autres
Publié: (2025)
Documents similaires
-
Tricky$^2$: Towards a Benchmark for Evaluating Human and LLM Error Interactions
par: Granger, Cole, et autres
Publié: (2026) -
A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm
par: Mastropaolo, Antonio, et autres
Publié: (2025) -
Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives
par: Khati, Dipin, et autres
Publié: (2025) -
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
par: Khati, Dipin, et autres
Publié: (2026) -
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
par: Palacio, David N., et autres
Publié: (2024)