Beyond Resolution Rates: Behavioral Drivers of Coding Agent Success and Failure
Fuente:
arXiv
Guardado en:
| Autores principales: | Mehtiyev, Tural, Assunção, Wesley |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories
por: Majgaonkar, Oorja, et al.
Publicado: (2025)
por: Majgaonkar, Oorja, et al.
Publicado: (2025)
Test Code Review in the Era of GitHub Actions: A Replication Study
por: Sun, Hui, et al.
Publicado: (2026)
por: Sun, Hui, et al.
Publicado: (2026)
Is LLM-Generated Code More Maintainable \& Reliable than Human-Written Code?
por: Molison, Alfred Santa, et al.
Publicado: (2025)
por: Molison, Alfred Santa, et al.
Publicado: (2025)
Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints
por: Li, Kefan, et al.
Publicado: (2026)
por: Li, Kefan, et al.
Publicado: (2026)
Code Quality Analysis of Translations from C to Rust
por: Tadesse, Biruk, et al.
Publicado: (2026)
por: Tadesse, Biruk, et al.
Publicado: (2026)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
por: Raghavendra, Mohit, et al.
Publicado: (2026)
por: Raghavendra, Mohit, et al.
Publicado: (2026)
Assessing the Bug-Proneness of Refactored Code: A Longitudinal Multi-Project Study
por: Ferreira, Isabella, et al.
Publicado: (2025)
por: Ferreira, Isabella, et al.
Publicado: (2025)
Where are the Hidden Gems? Applying Transformer Models for Design Discussion Detection
por: Arkoh, Lawrence, et al.
Publicado: (2026)
por: Arkoh, Lawrence, et al.
Publicado: (2026)
Beyond Translation Accuracy: Addressing False Failures in LLM-Based Code Translation
por: Rabbi, Fazle, et al.
Publicado: (2026)
por: Rabbi, Fazle, et al.
Publicado: (2026)
ABTest: Behavior-Driven Testing for AI Coding Agents
por: Dai, Wuyang, et al.
Publicado: (2026)
por: Dai, Wuyang, et al.
Publicado: (2026)
Model-based Maintenance and Evolution with GenAI: A Look into the Future
por: Marchezan, Luciano, et al.
Publicado: (2024)
por: Marchezan, Luciano, et al.
Publicado: (2024)
Contemporary Software Modernization: Perspectives and Challenges to Deal with Legacy Systems
por: Assunção, Wesley K. G., et al.
Publicado: (2024)
por: Assunção, Wesley K. G., et al.
Publicado: (2024)
Refactoring $\neq$ Bug-Inducing: Improving Defect Prediction with Code Change Tactics Analysis
por: Niu, Feifei, et al.
Publicado: (2025)
por: Niu, Feifei, et al.
Publicado: (2025)
Brevity is the Soul of Wit: Condensing Code Changes to Improve Commit Message Generation
por: Kuang, Hongyu, et al.
Publicado: (2025)
por: Kuang, Hongyu, et al.
Publicado: (2025)
Feature-oriented Test Case Selection and Prioritization During the Evolution of Highly-Configurable Systems
por: Mendonça, Willian D. F., et al.
Publicado: (2024)
por: Mendonça, Willian D. F., et al.
Publicado: (2024)
A Large-Scale Study on Developer Engagement and Expertise in Configurable Software System Projects
por: Milano, Karolina M., et al.
Publicado: (2025)
por: Milano, Karolina M., et al.
Publicado: (2025)
"Refactoring Runaway": Understanding and Mitigating Tangled Refactorings in Coding Agents for Issue Resolution
por: Tian, Zhao, et al.
Publicado: (2026)
por: Tian, Zhao, et al.
Publicado: (2026)
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
por: Guan, Hao, et al.
Publicado: (2026)
por: Guan, Hao, et al.
Publicado: (2026)
Beyond Language Barriers: Multi-Agent Coordination for Multi-Language Code Generation
por: Moumoula, Micheline Bénédicte, et al.
Publicado: (2025)
por: Moumoula, Micheline Bénédicte, et al.
Publicado: (2025)
Developer Experience with AI Coding Agents: HTTP Behavioral Signatures in Documentation Portals
por: Borysenko, Oleksii
Publicado: (2026)
por: Borysenko, Oleksii
Publicado: (2026)
Accountability in Code Review: The Role of Intrinsic Drivers and the Impact of LLMs
por: Alami, Adam, et al.
Publicado: (2025)
por: Alami, Adam, et al.
Publicado: (2025)
From PREVENTion to REACTion: Enhancing Failure Resolution in Naval Systems
por: Rossi, Maria Teresa, et al.
Publicado: (2025)
por: Rossi, Maria Teresa, et al.
Publicado: (2025)
On the Illusion of Success: An Empirical Study of Build Reruns and Silent Failures in Industrial CI
por: Aïdasso, Henri, et al.
Publicado: (2025)
por: Aïdasso, Henri, et al.
Publicado: (2025)
An Empirical Study on Challenges of Event Management in Microservice Architectures
por: Laigner, Rodrigo, et al.
Publicado: (2024)
por: Laigner, Rodrigo, et al.
Publicado: (2024)
CodeAgent: Autonomous Communicative Agents for Code Review
por: Tang, Xunzhu, et al.
Publicado: (2024)
por: Tang, Xunzhu, et al.
Publicado: (2024)
RefModel: Detecting Refactorings using Foundation Models
por: Simões, Pedro, et al.
Publicado: (2025)
por: Simões, Pedro, et al.
Publicado: (2025)
Making OpenAPI Documentation Agent-Ready: Detecting Documentation and REST Smells with a Multi-Agent LLM System
por: Lima, Rayfran Rocha, et al.
Publicado: (2026)
por: Lima, Rayfran Rocha, et al.
Publicado: (2026)
What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants
por: Hasan, Alif Al, et al.
Publicado: (2026)
por: Hasan, Alif Al, et al.
Publicado: (2026)
Relating Complexity, Explicitness, Effectiveness of Refactorings and Non-Functional Requirements: A Replication Study
por: Soares, Vinícius, et al.
Publicado: (2025)
por: Soares, Vinícius, et al.
Publicado: (2025)
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
por: Guo, Chengquan, et al.
Publicado: (2025)
por: Guo, Chengquan, et al.
Publicado: (2025)
Beyond Code Generation: Assessing Code LLM Maturity with Postconditions
por: He, Fusen, et al.
Publicado: (2024)
por: He, Fusen, et al.
Publicado: (2024)
Articulate but Wrong: Self-Review Failures in LLM-Based Code Modernization
por: Reddy, Gokul Chandra Purnachandra, et al.
Publicado: (2026)
por: Reddy, Gokul Chandra Purnachandra, et al.
Publicado: (2026)
Beyond Bug Fixes: An Empirical Investigation of Post-Merge Code Quality Issues in Agent-Generated Pull Requests
por: Cynthia, Shamse Tasnim, et al.
Publicado: (2026)
por: Cynthia, Shamse Tasnim, et al.
Publicado: (2026)
Compressing Code Context for LLM-based Issue Resolution
por: Jia, Haoxiang, et al.
Publicado: (2026)
por: Jia, Haoxiang, et al.
Publicado: (2026)
BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
por: Chen, Guoxin, et al.
Publicado: (2026)
por: Chen, Guoxin, et al.
Publicado: (2026)
Assessing, Exploiting, and Mitigating Syntactic Robustness Failures in LLM-Based Code Generation
por: Sarker, Laboni, et al.
Publicado: (2024)
por: Sarker, Laboni, et al.
Publicado: (2024)
Improving the Learning of Code Review Successive Tasks with Cross-Task Knowledge Distillation
por: Sghaier, Oussama Ben, et al.
Publicado: (2024)
por: Sghaier, Oussama Ben, et al.
Publicado: (2024)
CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
por: Zhang, Kechi, et al.
Publicado: (2024)
por: Zhang, Kechi, et al.
Publicado: (2024)
AgentFM: Role-Aware Failure Management for Distributed Databases with LLM-Driven Multi-Agents
por: Zhang, Lingzhe, et al.
Publicado: (2025)
por: Zhang, Lingzhe, et al.
Publicado: (2025)
Insights on Microservice Architecture Through the Eyes of Industry Practitioners
por: Nogueira, Vinicius L., et al.
Publicado: (2024)
por: Nogueira, Vinicius L., et al.
Publicado: (2024)
Ejemplares similares
-
Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories
por: Majgaonkar, Oorja, et al.
Publicado: (2025) -
Test Code Review in the Era of GitHub Actions: A Replication Study
por: Sun, Hui, et al.
Publicado: (2026) -
Is LLM-Generated Code More Maintainable \& Reliable than Human-Written Code?
por: Molison, Alfred Santa, et al.
Publicado: (2025) -
Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints
por: Li, Kefan, et al.
Publicado: (2026) -
Code Quality Analysis of Translations from C to Rust
por: Tadesse, Biruk, et al.
Publicado: (2026)