MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Jia, Jin, Deng, Zhiling, Chen, Zhuangbin, Wang, Yingqi, Zheng, Zibin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MicroRacer: Detecting Concurrency Bugs for Cloud Service Systems
by: Deng, Zhiling, et al.
Published: (2025)
by: Deng, Zhiling, et al.
Published: (2025)
A Survey on Failure Analysis and Fault Injection in AI Systems
by: Yu, Guangba, et al.
Published: (2024)
by: Yu, Guangba, et al.
Published: (2024)
Cast: Automated Resilience Testing for Production Cloud Service Systems
by: Chen, Zhuangbin, et al.
Published: (2026)
by: Chen, Zhuangbin, et al.
Published: (2026)
MAS-Algorithm: A Workflow for Solving Algorithmic Programming Problems with a Multi-Agent System
by: Xu, Yuliang, et al.
Published: (2026)
by: Xu, Yuliang, et al.
Published: (2026)
An Empirical Study of Agent Developer Practices in AI Agent Frameworks
by: Wang, Yanlin, et al.
Published: (2025)
by: Wang, Yanlin, et al.
Published: (2025)
Tracezip: Efficient Distributed Tracing via Trace Compression
by: Chen, Zhuangbin, et al.
Published: (2025)
by: Chen, Zhuangbin, et al.
Published: (2025)
AgentGit: A Version Control Framework for Reliable and Scalable LLM-Powered Multi-Agent Systems
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
A Self-Healing Framework for Reliable LLM-Based Autonomous Agents
by: Jeong, Cheonsu, et al.
Published: (2026)
by: Jeong, Cheonsu, et al.
Published: (2026)
ComBench: A Repo-level Real-world Benchmark for Compilation Error Repair
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024)
by: Zhang, Ziyao, et al.
Published: (2024)
FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning
by: Ding, Haoran, et al.
Published: (2026)
by: Ding, Haoran, et al.
Published: (2026)
RGD: Multi-LLM Based Agent Debugger via Refinement and Generation Guidance
by: Jin, Haolin, et al.
Published: (2024)
by: Jin, Haolin, et al.
Published: (2024)
GISclaw: A Comprehensive Open-Source LLM Agent System for Realistic Multi-Step Geospatial Analysis
by: Han, Jinzhen, et al.
Published: (2026)
by: Han, Jinzhen, et al.
Published: (2026)
On the Role of Fault Localization Context for LLM-Based Program Repair
by: Sepidband, Melika, et al.
Published: (2026)
by: Sepidband, Melika, et al.
Published: (2026)
Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults
by: Zhou, Zhenhao, et al.
Published: (2025)
by: Zhou, Zhenhao, et al.
Published: (2025)
An Empirical Evaluation of LLM-Based Approaches for Code Vulnerability Detection: RAG, SFT, and Dual-Agent Systems
by: Saju, Md Hasan, et al.
Published: (2026)
by: Saju, Md Hasan, et al.
Published: (2026)
MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution
by: Tao, Wei, et al.
Published: (2024)
by: Tao, Wei, et al.
Published: (2024)
From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems
by: Wang, Yawen, et al.
Published: (2026)
by: Wang, Yawen, et al.
Published: (2026)
LLM Collaboration With Multi-Agent Reinforcement Learning
by: Liu, Shuo, et al.
Published: (2025)
by: Liu, Shuo, et al.
Published: (2025)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
by: Huang, Donghao, et al.
Published: (2025)
by: Huang, Donghao, et al.
Published: (2025)
Efficient Failure Management for Multi-Agent Systems with Reasoning Trace Representation
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
by: Liang, Linxi, et al.
Published: (2025)
by: Liang, Linxi, et al.
Published: (2025)
LogPrism: Unifying Structure and Variable Encoding for Effective Log Compression
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
by: Ma, Ming, et al.
Published: (2025)
by: Ma, Ming, et al.
Published: (2025)
Metronome: Differentiated Delay Scheduling for Serverless Functions
by: Chen, Zhuangbin, et al.
Published: (2025)
by: Chen, Zhuangbin, et al.
Published: (2025)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
by: Guo, Lianghong, et al.
Published: (2025)
by: Guo, Lianghong, et al.
Published: (2025)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
by: Li, Yuanyang, et al.
Published: (2026)
by: Li, Yuanyang, et al.
Published: (2026)
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents
by: Hong, Sirui, et al.
Published: (2026)
by: Hong, Sirui, et al.
Published: (2026)
Adaptive Root Cause Localization for Microservice Systems with Multi-Agent Recursion-of-Thought
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
Identifying Performance-Sensitive Configurations in Software Systems through Code Analysis with LLM Agents
by: Wang, Zehao, et al.
Published: (2024)
by: Wang, Zehao, et al.
Published: (2024)
PyResBugs: A Dataset of Residual Python Bugs for Natural Language-Driven Fault Injection
by: Cotroneo, Domenico, et al.
Published: (2025)
by: Cotroneo, Domenico, et al.
Published: (2025)
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
by: Chen, Zhi, et al.
Published: (2026)
by: Chen, Zhi, et al.
Published: (2026)
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
by: He, Jiawei, et al.
Published: (2026)
by: He, Jiawei, et al.
Published: (2026)
The Dual-State Architecture for Reliable LLM Agents
by: Thompson, Matthew
Published: (2025)
by: Thompson, Matthew
Published: (2025)
TransAgent: Enhancing LLM-Based Code Translation via Fine-Grained Execution Alignment
by: Yuan, Zhiqiang, et al.
Published: (2024)
by: Yuan, Zhiqiang, et al.
Published: (2024)
Towards Reliable LLM-Driven Fuzz Testing: Vision and Road Ahead
by: Cheng, Yiran, et al.
Published: (2025)
by: Cheng, Yiran, et al.
Published: (2025)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
SAFEdit: Does Multi-Agent Decomposition Resolve the Reliability Challenges of Instructed Code Editing?
by: Tarshish, Noam, et al.
Published: (2026)
by: Tarshish, Noam, et al.
Published: (2026)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
by: Yan, Shuo, et al.
Published: (2025)
by: Yan, Shuo, et al.
Published: (2025)
Similar Items
-
MicroRacer: Detecting Concurrency Bugs for Cloud Service Systems
by: Deng, Zhiling, et al.
Published: (2025) -
A Survey on Failure Analysis and Fault Injection in AI Systems
by: Yu, Guangba, et al.
Published: (2024) -
Cast: Automated Resilience Testing for Production Cloud Service Systems
by: Chen, Zhuangbin, et al.
Published: (2026) -
MAS-Algorithm: A Workflow for Solving Algorithmic Programming Problems with a Multi-Agent System
by: Xu, Yuliang, et al.
Published: (2026) -
An Empirical Study of Agent Developer Practices in AI Agent Frameworks
by: Wang, Yanlin, et al.
Published: (2025)