An Empirical Study on Failures in Automated Issue Solving
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Simiao, Liu, Fang, Li, Liehao, Tan, Xin, Zhu, Yinghao, Lian, Xiaoli, Zhang, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SecureReviewer: Enhancing Large Language Models for Secure Code Review through Secure-aware Fine-tuning
by: Liu, Fang, et al.
Published: (2025)
by: Liu, Fang, et al.
Published: (2025)
MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repair
by: Liu, Simiao, et al.
Published: (2026)
by: Liu, Simiao, et al.
Published: (2026)
ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse
by: Liu, Simiao, et al.
Published: (2026)
by: Liu, Simiao, et al.
Published: (2026)
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
Enhancing Automated Program Repair with Solution Design
by: Zhao, Jiuang, et al.
Published: (2024)
by: Zhao, Jiuang, et al.
Published: (2024)
LLM-Driven Collaborative Model for Untangling Commits via Explicit and Implicit Dependency Reasoning
by: Hou, Bo, et al.
Published: (2025)
by: Hou, Bo, et al.
Published: (2025)
SEW: Self-Evolving Agentic Workflows for Automated Code Generation
by: Liu, Siwei, et al.
Published: (2025)
by: Liu, Siwei, et al.
Published: (2025)
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
by: Zan, Daoguang, et al.
Published: (2025)
by: Zan, Daoguang, et al.
Published: (2025)
CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language
by: Cheng, Junhang, et al.
Published: (2026)
by: Cheng, Junhang, et al.
Published: (2026)
Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines
by: Ning, Jingjie, et al.
Published: (2026)
by: Ning, Jingjie, et al.
Published: (2026)
DevEval: Evaluating Code Generation in Practical Software Projects
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Issue Localization via LLM-Driven Iterative Code Graph Searching
by: Jiang, Zhonghao, et al.
Published: (2025)
by: Jiang, Zhonghao, et al.
Published: (2025)
Towards Practical Defect-Focused Automated Code Review
by: Lu, Junyi, et al.
Published: (2025)
by: Lu, Junyi, et al.
Published: (2025)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
Your Simulation Runs but Solves the Wrong Physics: PDE-Grounded Intent Verification for LLM-Generated Multiphysics Simulation Code
by: Song, Zhenghan, et al.
Published: (2026)
by: Song, Zhenghan, et al.
Published: (2026)
AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications
by: Shen, Leming, et al.
Published: (2025)
by: Shen, Leming, et al.
Published: (2025)
Exploring Large Language Models in Resolving Environment-Related Crash Bugs: Localizing and Repairing
by: Du, Xueying, et al.
Published: (2023)
by: Du, Xueying, et al.
Published: (2023)
Multilingual Multimodal Software Developer for Code Generation
by: Chai, Linzheng, et al.
Published: (2025)
by: Chai, Linzheng, et al.
Published: (2025)
kRAIG: A Natural Language-Driven Agent for Automated DataOps Pipeline Generation
by: Siva, Rohan, et al.
Published: (2026)
by: Siva, Rohan, et al.
Published: (2026)
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
by: Luo, Jane, et al.
Published: (2025)
by: Luo, Jane, et al.
Published: (2025)
ChainStream: An LLM-based Framework for Unified Synthetic Sensing
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
Multi-agent Application System in Office Collaboration Scenarios
by: Sun, Songtao, et al.
Published: (2025)
by: Sun, Songtao, et al.
Published: (2025)
CodeV: Issue Resolving with Visual Data
by: Zhang, Linhao, et al.
Published: (2024)
by: Zhang, Linhao, et al.
Published: (2024)
Towards Automated Data Sciences with Natural Language and SageCopilot: Practices and Lessons Learned
by: Liao, Yuan, et al.
Published: (2024)
by: Liao, Yuan, et al.
Published: (2024)
WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models
by: Fan, Shengda, et al.
Published: (2024)
by: Fan, Shengda, et al.
Published: (2024)
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
by: Liu, Zuxin, et al.
Published: (2024)
by: Liu, Zuxin, et al.
Published: (2024)
PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
by: Chen, Simin, et al.
Published: (2024)
by: Chen, Simin, et al.
Published: (2024)
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
by: Cao, Yuhan, et al.
Published: (2025)
by: Cao, Yuhan, et al.
Published: (2025)
aiXcoder-7B: A Lightweight and Effective Large Language Model for Code Processing
by: Jiang, Siyuan, et al.
Published: (2024)
by: Jiang, Siyuan, et al.
Published: (2024)
Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
by: Kuang, Jiayi, et al.
Published: (2025)
by: Kuang, Jiayi, et al.
Published: (2025)
Terminal Agents Suffice for Enterprise Automation
by: Bechard, Patrice, et al.
Published: (2026)
by: Bechard, Patrice, et al.
Published: (2026)
Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
by: Liu, Fang, et al.
Published: (2024)
by: Liu, Fang, et al.
Published: (2024)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging
by: Islam, Md. Ashraful, et al.
Published: (2025)
by: Islam, Md. Ashraful, et al.
Published: (2025)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
by: Yan, Weixiang, et al.
Published: (2023)
by: Yan, Weixiang, et al.
Published: (2023)
Can ChatGPT replace StackOverflow? A Study on Robustness and Reliability of Large Language Model Code Generation
by: Zhong, Li, et al.
Published: (2023)
by: Zhong, Li, et al.
Published: (2023)
How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Code
by: Lee, Seonghyeon, et al.
Published: (2025)
by: Lee, Seonghyeon, et al.
Published: (2025)
AFlow: Automating Agentic Workflow Generation
by: Zhang, Jiayi, et al.
Published: (2024)
by: Zhang, Jiayi, et al.
Published: (2024)
Similar Items
-
SecureReviewer: Enhancing Large Language Models for Secure Code Review through Secure-aware Fine-tuning
by: Liu, Fang, et al.
Published: (2025) -
MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repair
by: Liu, Simiao, et al.
Published: (2026) -
ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse
by: Liu, Simiao, et al.
Published: (2026) -
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
by: Liu, Yi, et al.
Published: (2023) -
Enhancing Automated Program Repair with Solution Design
by: Zhao, Jiuang, et al.
Published: (2024)