Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Ruofan, Li, Yichen, Huo, Yintong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing LLM-Based Coding Tools through Native Integration of IDE-Derived Static Context
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
WebSuite: Systematically Evaluating Why Web Agents Fail
by: Li, Eric, et al.
Published: (2024)
by: Li, Eric, et al.
Published: (2024)
Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development
by: Wan, Yuxuan, et al.
Published: (2025)
by: Wan, Yuxuan, et al.
Published: (2025)
Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code
by: Kim, Myeongsoo, et al.
Published: (2026)
by: Kim, Myeongsoo, et al.
Published: (2026)
Next Edit Prediction: Learning to Predict Code Edits from Context and Interaction History
by: Lu, Ruofan, et al.
Published: (2025)
by: Lu, Ruofan, et al.
Published: (2025)
Why Attention Fails: A Taxonomy of Faults in Attention-Based Neural Networks
by: Jahan, Sigma, et al.
Published: (2025)
by: Jahan, Sigma, et al.
Published: (2025)
Automatically Generating UI Code from Screenshot: A Divide-and-Conquer-Based Approach
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
What Software Engineering Looks Like to AI Agents? -- An Empirical Study of AI-Only Technical Discourse on MoltBook
by: Huo, Junyu, et al.
Published: (2026)
by: Huo, Junyu, et al.
Published: (2026)
EfficientUICoder: Efficient MLLM-based UI Code Generation via Input and Output Token Compression
by: Xiao, Jingyu, et al.
Published: (2025)
by: Xiao, Jingyu, et al.
Published: (2025)
Five Fatal Assumptions: Why T-Shirt Sizing Systematically Fails for AI Projects
by: Soundaramourty, Raja, et al.
Published: (2026)
by: Soundaramourty, Raja, et al.
Published: (2026)
DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
by: Xiao, Jingyu, et al.
Published: (2025)
by: Xiao, Jingyu, et al.
Published: (2025)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026)
by: Ehsani, Ramtin, et al.
Published: (2026)
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
by: Xiong, Qian, et al.
Published: (2025)
by: Xiong, Qian, et al.
Published: (2025)
When Neural Code Completion Models Size up the Situation: Attaining Cheaper and Faster Completion through Dynamic Model Inference
by: Sun, Zhensu, et al.
Published: (2024)
by: Sun, Zhensu, et al.
Published: (2024)
End-to-End Automated Logging via Multi-Agent Framework
by: Zhong, Renyi, et al.
Published: (2025)
by: Zhong, Renyi, et al.
Published: (2025)
When the Specification Emerges: Benchmarking Faithfulness Loss in Long-Horizon Coding Agents
by: Yan, Lu, et al.
Published: (2026)
by: Yan, Lu, et al.
Published: (2026)
SpecAgent: A Speculative Retrieval and Forecasting Agent for Code Completion
by: Ma, George, et al.
Published: (2025)
by: Ma, George, et al.
Published: (2025)
When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning
by: Morabito, Roberto, et al.
Published: (2025)
by: Morabito, Roberto, et al.
Published: (2025)
MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
by: Yang, Ruofeng, et al.
Published: (2026)
by: Yang, Ruofeng, et al.
Published: (2026)
RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving
by: Wang, Huacan, et al.
Published: (2025)
by: Wang, Huacan, et al.
Published: (2025)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
by: Gandhi, Shubham, et al.
Published: (2025)
by: Gandhi, Shubham, et al.
Published: (2025)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion
by: Kotti, Zoe, et al.
Published: (2025)
by: Kotti, Zoe, et al.
Published: (2025)
Generating a Low-code Complete Workflow via Task Decomposition and RAG
by: Ayala, Orlando Marquez, et al.
Published: (2024)
by: Ayala, Orlando Marquez, et al.
Published: (2024)
Don't Complete It! Preventing Unhelpful Code Completion for Productive and Sustainable Neural Code Completion Systems
by: Sun, Zhensu, et al.
Published: (2022)
by: Sun, Zhensu, et al.
Published: (2022)
A Survey of Bugs in AI-Generated Code
by: Gao, Ruofan, et al.
Published: (2025)
by: Gao, Ruofan, et al.
Published: (2025)
A Self-Healing Framework for Reliable LLM-Based Autonomous Agents
by: Jeong, Cheonsu, et al.
Published: (2026)
by: Jeong, Cheonsu, et al.
Published: (2026)
AgentForge: Execution-Grounded Multi-Agent LLM Framework for Autonomous Software Engineering
by: Kumar, Rajesh, et al.
Published: (2026)
by: Kumar, Rajesh, et al.
Published: (2026)
Why Does My Transaction Fail? A First Look at Failed Transactions on the Solana Blockchain
by: Zheng, Xiaoye, et al.
Published: (2025)
by: Zheng, Xiaoye, et al.
Published: (2025)
Agyn: A Multi-Agent System for Team-Based Autonomous Software Engineering
by: Benkovich, Nikita, et al.
Published: (2026)
by: Benkovich, Nikita, et al.
Published: (2026)
RA-Gen: A Controllable Code Generation Framework Using ReAct for Multi-Agent Task Execution
by: Liu, Aofan, et al.
Published: (2025)
by: Liu, Aofan, et al.
Published: (2025)
HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale
by: Phan, Huy Nhat, et al.
Published: (2024)
by: Phan, Huy Nhat, et al.
Published: (2024)
MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
by: Fei, Xiang, et al.
Published: (2025)
by: Fei, Xiang, et al.
Published: (2025)
ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
by: Tawosi, Vali, et al.
Published: (2025)
by: Tawosi, Vali, et al.
Published: (2025)
Why Are AI Agent Involved Pull Requests (Fix-Related) Remain Unmerged? An Empirical Study
by: Alam, Khairul, et al.
Published: (2026)
by: Alam, Khairul, et al.
Published: (2026)
Tree-of-Code: A Tree-Structured Exploring Framework for End-to-End Code Generation and Execution in Complex Task Handling
by: Ni, Ziyi, et al.
Published: (2024)
by: Ni, Ziyi, et al.
Published: (2024)
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
by: Merrill, Mike A., et al.
Published: (2026)
by: Merrill, Mike A., et al.
Published: (2026)
When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
by: Larbi, Maya, et al.
Published: (2025)
by: Larbi, Maya, et al.
Published: (2025)
Similar Items
-
Enhancing LLM-Based Coding Tools through Native Integration of IDE-Derived Static Context
by: Li, Yichen, et al.
Published: (2024) -
WebSuite: Systematically Evaluating Why Web Agents Fail
by: Li, Eric, et al.
Published: (2024) -
Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development
by: Wan, Yuxuan, et al.
Published: (2025) -
Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code
by: Kim, Myeongsoo, et al.
Published: (2026) -
Next Edit Prediction: Learning to Predict Code Edits from Context and Interaction History
by: Lu, Ruofan, et al.
Published: (2025)