When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling
Fuente:
arXiv
Guardado en:
| Autores principales: | Islam, Niful, Ayon, Ragib Shahriar, Thomas, Deepak George, Ahmed, Shibbir, Wardat, Mohammad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SelfHeal: Empirical Fix Pattern Analysis and Bug Repair in LLM Agents
por: Islam, Niful, et al.
Publicado: (2026)
por: Islam, Niful, et al.
Publicado: (2026)
SpecPylot: Python Specification Generation using Large Language Models
por: Ayon, Ragib Shahariar, et al.
Publicado: (2026)
por: Ayon, Ragib Shahariar, et al.
Publicado: (2026)
AutoReSpec: A Framework for Generating Specification using Large Language Models
por: Ayon, Ragib Shahariar, et al.
Publicado: (2026)
por: Ayon, Ragib Shahariar, et al.
Publicado: (2026)
From Helpful to Trustworthy: LLM Agents for Pair Programming
por: Ayon, Ragib Shahariar
Publicado: (2026)
por: Ayon, Ragib Shahariar
Publicado: (2026)
Leveraging Data Characteristics for Bug Localization in Deep Learning Programs
por: Manke, Ruchira, et al.
Publicado: (2024)
por: Manke, Ruchira, et al.
Publicado: (2024)
An Empirical Study on LLM-based Agents for Automated Bug Fixing
por: Meng, Xiangxin, et al.
Publicado: (2024)
por: Meng, Xiangxin, et al.
Publicado: (2024)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
por: Zhu, Xinxue, et al.
Publicado: (2026)
por: Zhu, Xinxue, et al.
Publicado: (2026)
muPRL: A Mutation Testing Pipeline for Deep Reinforcement Learning based on Real Faults
por: Thomas, Deepak-George, et al.
Publicado: (2024)
por: Thomas, Deepak-George, et al.
Publicado: (2024)
Empirical Research on Utilizing LLM-based Agents for Automated Bug Fixing via LangGraph
por: Wang, Jialin, et al.
Publicado: (2025)
por: Wang, Jialin, et al.
Publicado: (2025)
AgentSZZ: Teaching the LLM Agent to Play Detective with Bug-Inducing Commits
por: Lyu, Yunbo, et al.
Publicado: (2026)
por: Lyu, Yunbo, et al.
Publicado: (2026)
MarsCode Agent: AI-native Automated Bug Fixing
por: Liu, Yizhou, et al.
Publicado: (2024)
por: Liu, Yizhou, et al.
Publicado: (2024)
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
por: Xiong, Qian, et al.
Publicado: (2025)
por: Xiong, Qian, et al.
Publicado: (2025)
Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in Deployment
por: Ahmed, Shibbir, et al.
Publicado: (2024)
por: Ahmed, Shibbir, et al.
Publicado: (2024)
BugGen: A Self-Correcting Multi-Agent LLM Pipeline for Realistic RTL Bug Synthesis
por: Jasper, Surya, et al.
Publicado: (2025)
por: Jasper, Surya, et al.
Publicado: (2025)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
por: Ehsani, Ramtin, et al.
Publicado: (2026)
por: Ehsani, Ramtin, et al.
Publicado: (2026)
LLM Agents for Automated Dependency Upgrades
por: Tawosi, Vali, et al.
Publicado: (2025)
por: Tawosi, Vali, et al.
Publicado: (2025)
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
por: Zheng, Mingwei, et al.
Publicado: (2025)
por: Zheng, Mingwei, et al.
Publicado: (2025)
AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents
por: Lin, Yixi, et al.
Publicado: (2026)
por: Lin, Yixi, et al.
Publicado: (2026)
Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
por: Lu, Ruofan, et al.
Publicado: (2025)
por: Lu, Ruofan, et al.
Publicado: (2025)
Mock Deep Testing: Toward Separate Development of Data and Models for Deep Learning
por: Manke, Ruchira, et al.
Publicado: (2025)
por: Manke, Ruchira, et al.
Publicado: (2025)
CMind: An AI Agent for Localizing C Memory Bugs
por: Su, Chia-Yi, et al.
Publicado: (2026)
por: Su, Chia-Yi, et al.
Publicado: (2026)
The Foundation Cracks: A Comprehensive Study on Bugs and Testing Practices in LLM Libraries
por: Jiang, Weipeng, et al.
Publicado: (2025)
por: Jiang, Weipeng, et al.
Publicado: (2025)
A Comprehensive Study on Automated Testing with the Software Lifecycle
por: Ali, Hussein Mohammed, et al.
Publicado: (2024)
por: Ali, Hussein Mohammed, et al.
Publicado: (2024)
A Comprehensive Study of Bugs in Modern Distributed Deep Learning Systems
por: Ma, Xiaoxue, et al.
Publicado: (2025)
por: Ma, Xiaoxue, et al.
Publicado: (2025)
Enhanced LLM-Based Framework for Predicting Null Pointer Dereference in Source Code
por: Sultan, Md. Fahim, et al.
Publicado: (2024)
por: Sultan, Md. Fahim, et al.
Publicado: (2024)
How and Why Agents Can Identify Bug-Introducing Commits
por: Risse, Niklas, et al.
Publicado: (2026)
por: Risse, Niklas, et al.
Publicado: (2026)
Investigating the Impact of Code Comment Inconsistency on Bug Introducing
por: Radmanesh, Shiva, et al.
Publicado: (2024)
por: Radmanesh, Shiva, et al.
Publicado: (2024)
An Empirical Study on Leveraging Images in Automated Bug Report Reproduction
por: Wang, Dingbang, et al.
Publicado: (2025)
por: Wang, Dingbang, et al.
Publicado: (2025)
Evaluating LLM Agents on Automated Software Analysis Tasks
por: Bouzenia, Islem, et al.
Publicado: (2026)
por: Bouzenia, Islem, et al.
Publicado: (2026)
LogiAgent: Automated Logical Testing for REST Systems with LLM-Based Multi-Agents
por: Zhang, Ke, et al.
Publicado: (2025)
por: Zhang, Ke, et al.
Publicado: (2025)
Leveraging LLM Agents for Automated Video Game Testing
por: Wang, Chengjia, et al.
Publicado: (2025)
por: Wang, Chengjia, et al.
Publicado: (2025)
Characterizing Real-World Bugs in Tile Programs for Automated Bug Detection
por: Rathnasuriya, Ravishka, et al.
Publicado: (2026)
por: Rathnasuriya, Ravishka, et al.
Publicado: (2026)
LLM-Based Detection of Tangled Code Changes for Higher-Quality Method-Level Bug Datasets
por: Opu, Md Nahidul Islam, et al.
Publicado: (2025)
por: Opu, Md Nahidul Islam, et al.
Publicado: (2025)
A Multi-Agent Framework for Automated Exploit Generation with Constraint-Guided Comprehension and Reflection
por: Chen, Siyi, et al.
Publicado: (2026)
por: Chen, Siyi, et al.
Publicado: (2026)
From REST to MCP: An Empirical Study of API Wrapping and Automated Server Generation for LLM Agents
por: Mastouri, Meriem, et al.
Publicado: (2025)
por: Mastouri, Meriem, et al.
Publicado: (2025)
Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents
por: Liu, Xiang, et al.
Publicado: (2026)
por: Liu, Xiang, et al.
Publicado: (2026)
A Comprehensive Study of Bug-Fix Patterns in Autonomous Driving Systems
por: Chen, Yuntianyi, et al.
Publicado: (2025)
por: Chen, Yuntianyi, et al.
Publicado: (2025)
BugsRepo: A Comprehensive Curated Dataset of Bug Reports, Comments and Contributors Information from Bugzilla
por: Acharya, Jagrit, et al.
Publicado: (2025)
por: Acharya, Jagrit, et al.
Publicado: (2025)
PromptDebt: A Comprehensive Study of Technical Debt Across LLM Projects
por: Aljohani, Ahmed, et al.
Publicado: (2025)
por: Aljohani, Ahmed, et al.
Publicado: (2025)
SysPro: Reproducing System-level Concurrency Bugs from Bug Reports
por: Zaman, Tarannum Shaila, et al.
Publicado: (2026)
por: Zaman, Tarannum Shaila, et al.
Publicado: (2026)
Ejemplares similares
-
SelfHeal: Empirical Fix Pattern Analysis and Bug Repair in LLM Agents
por: Islam, Niful, et al.
Publicado: (2026) -
SpecPylot: Python Specification Generation using Large Language Models
por: Ayon, Ragib Shahariar, et al.
Publicado: (2026) -
AutoReSpec: A Framework for Generating Specification using Large Language Models
por: Ayon, Ragib Shahariar, et al.
Publicado: (2026) -
From Helpful to Trustworthy: LLM Agents for Pair Programming
por: Ayon, Ragib Shahariar
Publicado: (2026) -
Leveraging Data Characteristics for Bug Localization in Deep Learning Programs
por: Manke, Ruchira, et al.
Publicado: (2024)