Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Jingyue, Storhaug, André |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study
par: Storhaug, André, et autres
Publié: (2024)
par: Storhaug, André, et autres
Publié: (2024)
Favia: Forensic Agent for Vulnerability-fix Identification and Analysis
par: Storhaug, André, et autres
Publié: (2026)
par: Storhaug, André, et autres
Publié: (2026)
Agentic AI Software Engineers: Programming with Trust
par: Roychoudhury, Abhik, et autres
Publié: (2025)
par: Roychoudhury, Abhik, et autres
Publié: (2025)
Agentic AI for Software: thoughts from Software Engineering community
par: Roychoudhury, Abhik
Publié: (2025)
par: Roychoudhury, Abhik
Publié: (2025)
Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies
par: Angermeir, Florian, et autres
Publié: (2025)
par: Angermeir, Florian, et autres
Publié: (2025)
Agentic Software Engineering: Foundational Pillars and a Research Roadmap
par: Hassan, Ahmed E., et autres
Publié: (2025)
par: Hassan, Ahmed E., et autres
Publié: (2025)
Unified Software Engineering Agent as AI Software Engineer
par: Applis, Leonhard, et autres
Publié: (2025)
par: Applis, Leonhard, et autres
Publié: (2025)
On the Replicability and Reproducibility of Deep Learning in Software Engineering
par: Liu, Chao, et autres
Publié: (2020)
par: Liu, Chao, et autres
Publié: (2020)
LLM-Based Agentic Systems for Software Engineering: Challenges and Opportunities
par: Tang, Yongjian, et autres
Publié: (2026)
par: Tang, Yongjian, et autres
Publié: (2026)
From Code-Centric to Intent-Centric Software Engineering: A Reflexive Thematic Analysis of Generative AI, Agentic Systems, and Engineering Accountability
par: De La Cruz, Elyson
Publié: (2026)
par: De La Cruz, Elyson
Publié: (2026)
SBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration, and Reproducibility Evaluation
par: Radanliev, Petar, et autres
Publié: (2026)
par: Radanliev, Petar, et autres
Publié: (2026)
The Semi-Executable Stack: Agentic Software Engineering and the Expanding Scope of SE
par: Feldt, Robert, et autres
Publié: (2026)
par: Feldt, Robert, et autres
Publié: (2026)
AI-Tutoring in Software Engineering Education
par: Frankford, Eduard, et autres
Publié: (2024)
par: Frankford, Eduard, et autres
Publié: (2024)
Beyond Human-Readable: Rethinking Software Engineering Conventions for the Agentic Development Era
par: Ustynov, Dmytro
Publié: (2026)
par: Ustynov, Dmytro
Publié: (2026)
A Systematic Literature Review on Explainability for Machine/Deep Learning-based Software Engineering Research
par: Cao, Sicong, et autres
Publié: (2024)
par: Cao, Sicong, et autres
Publié: (2024)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
par: Fan, Zhiyu, et autres
Publié: (2025)
par: Fan, Zhiyu, et autres
Publié: (2025)
Agentic AI in 6G Software Businesses: A Layered Maturity Model
par: Zohaib, Muhammad, et autres
Publié: (2025)
par: Zohaib, Muhammad, et autres
Publié: (2025)
Software Reuse in the Generative AI Era: From Cargo Cult Towards AI Native Software Engineering
par: Mikkonen, Tommi, et autres
Publié: (2025)
par: Mikkonen, Tommi, et autres
Publié: (2025)
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
par: Xu, Xinbo, et autres
Publié: (2026)
par: Xu, Xinbo, et autres
Publié: (2026)
Beyond the 'Diff': Addressing Agentic Entropy in Agentic Software Development
par: Casserini, Matteo, et autres
Publié: (2026)
par: Casserini, Matteo, et autres
Publié: (2026)
Will AI replace Software Engineers? Do not hold your breath
par: Roychoudhury, Abhik, et autres
Publié: (2025)
par: Roychoudhury, Abhik, et autres
Publié: (2025)
Generative AI and Empirical Software Engineering: A Paradigm Shift
par: Treude, Christoph, et autres
Publié: (2025)
par: Treude, Christoph, et autres
Publié: (2025)
ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation
par: Xianpeng, et autres
Publié: (2026)
par: Xianpeng, et autres
Publié: (2026)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
par: Fandina, Ora Nova, et autres
Publié: (2025)
par: Fandina, Ora Nova, et autres
Publié: (2025)
Shift-Up: A Framework for Software Engineering Guardrails in AI-native Software Development -- Initial Findings
par: Lipsanen, Petrus, et autres
Publié: (2026)
par: Lipsanen, Petrus, et autres
Publié: (2026)
Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering
par: Salim, Mohamad, et autres
Publié: (2026)
par: Salim, Mohamad, et autres
Publié: (2026)
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
par: Adamenko, Pavel, et autres
Publié: (2025)
par: Adamenko, Pavel, et autres
Publié: (2025)
Challenges and Paths Towards AI for Software Engineering
par: Gu, Alex, et autres
Publié: (2025)
par: Gu, Alex, et autres
Publié: (2025)
Mitigating Prompt-Induced Cognitive Biases in General-Purpose AI for Software Engineering
par: Sovrano, Francesco, et autres
Publié: (2026)
par: Sovrano, Francesco, et autres
Publié: (2026)
AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents
par: Zhong, Hailin, et autres
Publié: (2026)
par: Zhong, Hailin, et autres
Publié: (2026)
Evidence-Driven Decision Support for AI Model Selection in Research Software Engineering
par: Joonbakhsh, Alireza, et autres
Publié: (2025)
par: Joonbakhsh, Alireza, et autres
Publié: (2025)
A Gray Literature Study on Fairness Requirements in AI-enabled Software Engineering
par: Nguyen, Thanh, et autres
Publié: (2025)
par: Nguyen, Thanh, et autres
Publié: (2025)
Greening AI-enabled Systems with Software Engineering: A Research Agenda for Environmentally Sustainable AI Practices
par: Cruz, Luís, et autres
Publié: (2025)
par: Cruz, Luís, et autres
Publié: (2025)
When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs'' for Human-AI Interaction
par: Xing, Zhenchang, et autres
Publié: (2025)
par: Xing, Zhenchang, et autres
Publié: (2025)
Test Before You Deploy: Governing Updates in the LLM Supply Chain
par: Chishti, Mohd Sameen, et autres
Publié: (2026)
par: Chishti, Mohd Sameen, et autres
Publié: (2026)
EvoClaw: Evaluating AI Agents on Continuous Software Evolution
par: Deng, Gangda, et autres
Publié: (2026)
par: Deng, Gangda, et autres
Publié: (2026)
LLMs: A Game-Changer for Software Engineers?
par: Haque, Md Asraful
Publié: (2024)
par: Haque, Md Asraful
Publié: (2024)
Software Performance Engineering for Foundation Model-Powered Software
par: Zhang, Haoxiang, et autres
Publié: (2024)
par: Zhang, Haoxiang, et autres
Publié: (2024)
An LLM Agentic Approach for Legal-Critical Software: A Case Study for Tax Prep Software
par: Gogani-Khiabani, Sina, et autres
Publié: (2025)
par: Gogani-Khiabani, Sina, et autres
Publié: (2025)
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
par: Trinkenreich, Bianca, et autres
Publié: (2026)
par: Trinkenreich, Bianca, et autres
Publié: (2026)
Documents similaires
-
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study
par: Storhaug, André, et autres
Publié: (2024) -
Favia: Forensic Agent for Vulnerability-fix Identification and Analysis
par: Storhaug, André, et autres
Publié: (2026) -
Agentic AI Software Engineers: Programming with Trust
par: Roychoudhury, Abhik, et autres
Publié: (2025) -
Agentic AI for Software: thoughts from Software Engineering community
par: Roychoudhury, Abhik
Publié: (2025) -
Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies
par: Angermeir, Florian, et autres
Publié: (2025)