A Comprehensive Empirical Evaluation of Agent Frameworks on Code-centric Software Engineering Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Zhuowen, Gao, Cuifeng, Fan, Chunsong, Yang, Wenzhang, Xue, Yinxing, Zhang, Lijun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scheduzz: Constraint-based Fuzz Driver Generation with Dual Scheduling
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
An Empirical Study of Speculative Decoding on Software Engineering Tasks
by: Li, Yijia, et al.
Published: (2026)
by: Li, Yijia, et al.
Published: (2026)
TGMM: Combining Parse Tree with GPU for Scalable Multilingual and Multi-Granularity Code Clone Detection
by: Ye, Yuhang, et al.
Published: (2024)
by: Ye, Yuhang, et al.
Published: (2024)
Beyond Isolated Tasks: A Framework for Evaluating Coding Agents on Sequential Software Evolution
by: Shastry, KN Ajay, et al.
Published: (2026)
by: Shastry, KN Ajay, et al.
Published: (2026)
Strengthening Programming Comprehension in Large Language Models through Code Generation
by: Ren, Xiaoning, et al.
Published: (2025)
by: Ren, Xiaoning, et al.
Published: (2025)
Chain of Draft for Software Engineering: Challenges in Applying Concise Reasoning to Code Tasks
by: Yang, Shaoyi
Published: (2025)
by: Yang, Shaoyi
Published: (2025)
HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale
by: Phan, Huy Nhat, et al.
Published: (2024)
by: Phan, Huy Nhat, et al.
Published: (2024)
PAGENT: Learning to Patch Software Engineering Agents
by: Xue, Haoran, et al.
Published: (2025)
by: Xue, Haoran, et al.
Published: (2025)
Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation
by: Xiong, Qian, et al.
Published: (2025)
by: Xiong, Qian, et al.
Published: (2025)
A Framework for Using LLMs for Repository Mining Studies in Empirical Software Engineering
by: de Martino, Vincenzo, et al.
Published: (2024)
by: de Martino, Vincenzo, et al.
Published: (2024)
Not All RAGs Are Created Equal: A Component-Wise Empirical Study for Software Engineering Tasks
by: Ke, Qiang, et al.
Published: (2026)
by: Ke, Qiang, et al.
Published: (2026)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Evaluating LLM Agents on Automated Software Analysis Tasks
by: Bouzenia, Islem, et al.
Published: (2026)
by: Bouzenia, Islem, et al.
Published: (2026)
Code Comments for Quantum Software Development Kits: An Empirical Study on Qiskit
by: Zhou, Zenghui, et al.
Published: (2025)
by: Zhou, Zenghui, et al.
Published: (2025)
Towards Causal Analysis of Empirical Software Engineering Data: The Impact of Programming Languages on Coding Competitions
by: Furia, Carlo A., et al.
Published: (2023)
by: Furia, Carlo A., et al.
Published: (2023)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
by: Sonwane, Atharv, et al.
Published: (2026)
by: Sonwane, Atharv, et al.
Published: (2026)
Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks
by: Hu, Xing, et al.
Published: (2025)
by: Hu, Xing, et al.
Published: (2025)
An Empirical Investigation of the Experiences of Dyslexic Software Engineers
by: Cruz, Marcos Vinicius, et al.
Published: (2025)
by: Cruz, Marcos Vinicius, et al.
Published: (2025)
How Natural Language Proficiency Shapes GenAI Code for Software Engineering Tasks
by: Rojpaisarnkit, Ruksit, et al.
Published: (2025)
by: Rojpaisarnkit, Ruksit, et al.
Published: (2025)
Empirical Studies on Quantum Optimization for Software Engineering: A Systematic Analysis
by: Zhang, Man, et al.
Published: (2025)
by: Zhang, Man, et al.
Published: (2025)
An Empirical Study of Knowledge Distillation for Code Understanding Tasks
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Which Prompting Technique Should I Use? An Empirical Investigation of Prompting Techniques for Software Engineering Tasks
by: Santana Jr, E. G., et al.
Published: (2025)
by: Santana Jr, E. G., et al.
Published: (2025)
Applying Bayesian Analysis Guidelines to Empirical Software Engineering Data: The Case of Programming Languages and Code Quality
by: Furia, Carlo A., et al.
Published: (2021)
by: Furia, Carlo A., et al.
Published: (2021)
IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks
by: Mateega, Spencer, et al.
Published: (2026)
by: Mateega, Spencer, et al.
Published: (2026)
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
by: Badertdinov, Ibragim, et al.
Published: (2025)
by: Badertdinov, Ibragim, et al.
Published: (2025)
Automated Code-centric Software Vulnerability Assessment: How Far Are We? An Empirical Study in C/C++
by: Nguyen, Anh The, et al.
Published: (2024)
by: Nguyen, Anh The, et al.
Published: (2024)
Integrating Positionality Statements in Empirical Software Engineering Research
by: de Sousa, Breno Felix, et al.
Published: (2024)
by: de Sousa, Breno Felix, et al.
Published: (2024)
An Empirical Study of Generative AI Adoption in Software Engineering
by: Giray, Görkem, et al.
Published: (2025)
by: Giray, Görkem, et al.
Published: (2025)
Mitigating Omitted Variable Bias in Empirical Software Engineering
by: Furia, Carlo A., et al.
Published: (2025)
by: Furia, Carlo A., et al.
Published: (2025)
SEAlign: Alignment Training for Software Engineering Agent
by: Zhang, Kechi, et al.
Published: (2025)
by: Zhang, Kechi, et al.
Published: (2025)
Exploring the Power of Diffusion Large Language Models for Software Engineering: An Empirical Investigation
by: Zhang, Jingyao, et al.
Published: (2025)
by: Zhang, Jingyao, et al.
Published: (2025)
Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action
by: Pujar, Saurabh, et al.
Published: (2025)
by: Pujar, Saurabh, et al.
Published: (2025)
Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
by: Zhang, Ruixin, et al.
Published: (2026)
by: Zhang, Ruixin, et al.
Published: (2026)
Agentic AI in the Software Development Lifecycle: Architecture, Empirical Evidence, and the Reshaping of Software Engineering
by: Bhati, Happy
Published: (2026)
by: Bhati, Happy
Published: (2026)
Reward Engineering for Reinforcement Learning in Software Tasks
by: Masud, Md Rayhanul, et al.
Published: (2026)
by: Masud, Md Rayhanul, et al.
Published: (2026)
Unified Software Engineering Agent as AI Software Engineer
by: Applis, Leonhard, et al.
Published: (2025)
by: Applis, Leonhard, et al.
Published: (2025)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
by: Liu, Shuhan, et al.
Published: (2026)
by: Liu, Shuhan, et al.
Published: (2026)
Teaching Empirical Research Methods in Software Engineering: An Editorial Introduction
by: Mendez, Daniel, et al.
Published: (2025)
by: Mendez, Daniel, et al.
Published: (2025)
Designing a Syllabus for a Course on Empirical Software Engineering
by: Avgeriou, Paris, et al.
Published: (2025)
by: Avgeriou, Paris, et al.
Published: (2025)
Teaching Simulation as a Research Method in Empirical Software Engineering
by: de França, Breno Bernard Nicolau, et al.
Published: (2025)
by: de França, Breno Bernard Nicolau, et al.
Published: (2025)
Similar Items
-
Scheduzz: Constraint-based Fuzz Driver Generation with Dual Scheduling
by: Li, Yan, et al.
Published: (2025) -
An Empirical Study of Speculative Decoding on Software Engineering Tasks
by: Li, Yijia, et al.
Published: (2026) -
TGMM: Combining Parse Tree with GPU for Scalable Multilingual and Multi-Granularity Code Clone Detection
by: Ye, Yuhang, et al.
Published: (2024) -
Beyond Isolated Tasks: A Framework for Evaluating Coding Agents on Sequential Software Evolution
by: Shastry, KN Ajay, et al.
Published: (2026) -
Strengthening Programming Comprehension in Large Language Models through Code Generation
by: Ren, Xiaoning, et al.
Published: (2025)