Dissecting Bug Triggers and Failure Modes in Modern Agentic Frameworks: An Empirical Study
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xiaowen, Zhang, Hannuo, Tan, Shin Hwei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Study of Refactoring Engine Bugs
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
Investigating Code Reuse in Software Redesign: A Case Study
by: Zhang, Xiaowen, et al.
Published: (2026)
by: Zhang, Xiaowen, et al.
Published: (2026)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
by: Zhu, Xinxue, et al.
Published: (2026)
by: Zhu, Xinxue, et al.
Published: (2026)
Testing Refactoring Engine via Historical Bug Report driven LLM
by: Wang, Haibo, et al.
Published: (2025)
by: Wang, Haibo, et al.
Published: (2025)
An Empirical Study of False Negatives and Positives of Static Code Analyzers From the Perspective of Historical Issues
by: Cui, Han, et al.
Published: (2024)
by: Cui, Han, et al.
Published: (2024)
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
by: Tan, Honghao, et al.
Published: (2026)
by: Tan, Honghao, et al.
Published: (2026)
LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code
by: Guo, Liwei, et al.
Published: (2025)
by: Guo, Liwei, et al.
Published: (2025)
Understanding and Detecting Annotation-Induced Faults of Static Analyzers
by: Zhang, Huaien, et al.
Published: (2024)
by: Zhang, Huaien, et al.
Published: (2024)
Guiding ChatGPT to Fix Web UI Tests via Explanation-Consistency Checking
by: Xu, Zhuolin, et al.
Published: (2023)
by: Xu, Zhuolin, et al.
Published: (2023)
LLM-Guided Issue Generation from Uncovered Code Segments
by: Pressato, Diany, et al.
Published: (2026)
by: Pressato, Diany, et al.
Published: (2026)
An Empirical Study on Leveraging Images in Automated Bug Report Reproduction
by: Wang, Dingbang, et al.
Published: (2025)
by: Wang, Dingbang, et al.
Published: (2025)
An Empirical Study on the Capability of LLMs in Decomposing Bug Reports
by: Chen, Zhiyuan, et al.
Published: (2025)
by: Chen, Zhiyuan, et al.
Published: (2025)
Exploring the Jupyter Ecosystem: An Empirical Study of Bugs and Vulnerabilities
by: Jiang, Wenyuan, et al.
Published: (2025)
by: Jiang, Wenyuan, et al.
Published: (2025)
Bug Priority Change: An Empirical Study on Apache Projects
by: Li, Zengyang, et al.
Published: (2024)
by: Li, Zengyang, et al.
Published: (2024)
An Empirical Study on the Classification of Bug Reports with Machine Learning
by: Andrade, Renato, et al.
Published: (2025)
by: Andrade, Renato, et al.
Published: (2025)
An Empirical Study of Interaction Bugs in ROS-based Software
by: Chen, Zhixiang, et al.
Published: (2025)
by: Chen, Zhixiang, et al.
Published: (2025)
Ethics Testing: Proactive Identification of Generative AI System Harms
by: Tan, Shin Hwei, et al.
Published: (2026)
by: Tan, Shin Hwei, et al.
Published: (2026)
Automated Harmfulness Testing for Code Large Language Models
by: Tan, Honghao, et al.
Published: (2025)
by: Tan, Honghao, et al.
Published: (2025)
Understanding Bug-Reproducing Tests: A First Empirical Study
by: Hora, Andre, et al.
Published: (2026)
by: Hora, Andre, et al.
Published: (2026)
Characterizing Bugs in Login Processes of Android Applications: An Empirical Study
by: Zhou, Zixu, et al.
Published: (2025)
by: Zhou, Zixu, et al.
Published: (2025)
BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
by: Ma, Yinghang, et al.
Published: (2025)
by: Ma, Yinghang, et al.
Published: (2025)
A Comprehensive Study of Bugs in Modern Distributed Deep Learning Systems
by: Ma, Xiaoxue, et al.
Published: (2025)
by: Ma, Xiaoxue, et al.
Published: (2025)
An Empirical Evaluation of Modern MLOps Frameworks
by: Marcos-Mercadé, Jon, et al.
Published: (2026)
by: Marcos-Mercadé, Jon, et al.
Published: (2026)
From Logic to Toolchains: An Empirical Study of Bugs in the TypeScript Ecosystem
by: Tang, TianYi, et al.
Published: (2026)
by: Tang, TianYi, et al.
Published: (2026)
Does Programming Language Matter? An Empirical Study of Fuzzing Bug Detection
by: Shirai, Tatsuya, et al.
Published: (2026)
by: Shirai, Tatsuya, et al.
Published: (2026)
An Empirical Study on Embodied Artificial Intelligence Robot (EAIR) Software Bugs
by: Liao, Zeqin, et al.
Published: (2025)
by: Liao, Zeqin, et al.
Published: (2025)
Moving beyond Deletions: Program Simplification via Diverse Program Transformations
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
Agentic Frameworks for Reasoning Tasks: An Empirical Study
by: Rasheed, Zeeshan, et al.
Published: (2026)
by: Rasheed, Zeeshan, et al.
Published: (2026)
Understanding Bugs in Quantum Simulators: An Empirical Study
by: Upadhyay, Krishna, et al.
Published: (2026)
by: Upadhyay, Krishna, et al.
Published: (2026)
Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
by: Zhang, Ruixin, et al.
Published: (2026)
by: Zhang, Ruixin, et al.
Published: (2026)
An Empirical Study on the Characteristics of Database Access Bugs in Java Applications
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Are Benchmark Tests Strong Enough? Mutation-Guided Diagnosis and Augmentation of Regression Suites
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
From Historical Patches to Repair Plans: Outcome-Conditioned Reasoning for Repository-Level Program Repair
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
What Makes Code Generation Ethically Sourced?
by: Xu, Zhuolin, et al.
Published: (2025)
by: Xu, Zhuolin, et al.
Published: (2025)
Tumbling Down the Rabbit Hole: How do Assisting Exploration Strategies Facilitate Grey-box Fuzzing?
by: Wu, Mingyuan, et al.
Published: (2024)
by: Wu, Mingyuan, et al.
Published: (2024)
AgenticSZZ: Temporal Knowledge Graph-Guided Agentic Bug-Inducing Commit Identification
by: Shi, Yu, et al.
Published: (2026)
by: Shi, Yu, et al.
Published: (2026)
An Empirical Study on Bug Severity Estimation using Source Code Metrics and Static Analysis
by: Mashhadi, Ehsan, et al.
Published: (2022)
by: Mashhadi, Ehsan, et al.
Published: (2022)
Characterizing Bugs and Quality Attributes in Quantum Software: A Large-Scale Empirical Study
by: Yousuf, Mir Mohammad, et al.
Published: (2025)
by: Yousuf, Mir Mohammad, et al.
Published: (2025)
An Empirical Study on Failures in Automated Issue Solving
by: Liu, Simiao, et al.
Published: (2025)
by: Liu, Simiao, et al.
Published: (2025)
BugScope: Learn to Find Bugs Like Human
by: Guo, Jinyao, et al.
Published: (2025)
by: Guo, Jinyao, et al.
Published: (2025)
Similar Items
-
An Empirical Study of Refactoring Engine Bugs
by: Wang, Haibo, et al.
Published: (2024) -
Investigating Code Reuse in Software Redesign: A Case Study
by: Zhang, Xiaowen, et al.
Published: (2026) -
An Empirical Study of Bugs in Modern LLM Agent Frameworks
by: Zhu, Xinxue, et al.
Published: (2026) -
Testing Refactoring Engine via Historical Bug Report driven LLM
by: Wang, Haibo, et al.
Published: (2025) -
An Empirical Study of False Negatives and Positives of Static Code Analyzers From the Perspective of Historical Issues
by: Cui, Han, et al.
Published: (2024)