AgentFM: Role-Aware Failure Management for Distributed Databases with LLM-Driven Multi-Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Lingzhe, Zhai, Yunpeng, Jia, Tong, Huang, Xiaosong, Duan, Chiming, Li, Ying |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
LogDB: Multivariate Log-based Failure Diagnosis for Distributed Databases (Extended from MultiLog)
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
ThinkFL: Self-Refining Failure Localization for Microservice Systems via Reinforcement Fine-Tuning
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
Efficient Failure Management for Multi-Agent Systems with Reasoning Trace Representation
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
Adaptive Root Cause Localization for Microservice Systems with Multi-Agent Recursion-of-Thought
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
Agentic Memory Enhanced Recursive Reasoning for Root Cause Localization in Microservices
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
E2E-REME: Towards End-to-End Microservices Auto-Remediation via Experience-Simulation Reinforcement Fine-Tuning
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
Hypothesize-Then-Verify: Speculative Root Cause Analysis for Microservices with Pathwise Parallelism
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
Multivariate Log-based Anomaly Detection for Distributed Database
by: Zhang, Lingzhe, et al.
Published: (2024)
by: Zhang, Lingzhe, et al.
Published: (2024)
Towards In-Depth Root Cause Localization for Microservices with Multi-Agent Recursion-of-Thought
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
RuntimeSlicer: Towards Generalizable Unified Runtime State Representation for Failure Management
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
Walk the Talk: Is Your Log-based Software Reliability Maintenance System Really Reliable?
by: He, Minghua, et al.
Published: (2025)
by: He, Minghua, et al.
Published: (2025)
MicroRemed: Benchmarking LLMs in Microservices Remediation
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
United We Stand: Towards End-to-End Log-based Fault Diagnosis via Interactive Multi-Task Learning
by: He, Minghua, et al.
Published: (2025)
by: He, Minghua, et al.
Published: (2025)
A Survey of AIOps for Failure Management in the Era of Large Language Models
by: Zhang, Lingzhe, et al.
Published: (2024)
by: Zhang, Lingzhe, et al.
Published: (2024)
Enhancing Web Service Anomaly Detection via Fine-grained Multi-modal Association and Frequency Domain Analysis
by: Yang, Xixuan, et al.
Published: (2025)
by: Yang, Xixuan, et al.
Published: (2025)
Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive Filtering
by: Xiong, Yunpeng, et al.
Published: (2026)
by: Xiong, Yunpeng, et al.
Published: (2026)
Managing Uncertainty in LLM-based Multi-Agent System Operation
by: Zhang, Man, et al.
Published: (2026)
by: Zhang, Man, et al.
Published: (2026)
RepoTransAgent: Multi-Agent LLM Framework for Repository-Aware Code Translation
by: Guan, Ziqi, et al.
Published: (2025)
by: Guan, Ziqi, et al.
Published: (2025)
Reducing Events to Augment Log-based Anomaly Detection Models: An Empirical Study
by: Zhang, Lingzhe, et al.
Published: (2024)
by: Zhang, Lingzhe, et al.
Published: (2024)
From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
A Practical Framework for Flaky Failure Triage in Distributed Database Continuous Integration
by: Zhu, Jun-Peng, et al.
Published: (2026)
by: Zhu, Jun-Peng, et al.
Published: (2026)
LogiAgent: Automated Logical Testing for REST Systems with LLM-Based Multi-Agents
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
Atomizer: An LLM-based Collaborative Multi-Agent Framework for Intent-Driven Commit Untangling
by: Zhu, Kangchen, et al.
Published: (2026)
by: Zhu, Kangchen, et al.
Published: (2026)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
by: Zhu, Xinxue, et al.
Published: (2026)
by: Zhu, Xinxue, et al.
Published: (2026)
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
by: Ma, Ming, et al.
Published: (2025)
by: Ma, Ming, et al.
Published: (2025)
Traceability and Accountability in Role-Specialized Multi-Agent LLM Pipelines
by: Barrak, Amine
Published: (2025)
by: Barrak, Amine
Published: (2025)
REAgent: Requirement-Driven LLM Agents for Software Issue Resolution
by: Kuang, Shiqi, et al.
Published: (2026)
by: Kuang, Shiqi, et al.
Published: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
GraphCodeAgent: Dual Graph-Guided LLM Agent for Retrieval-Augmented Repo-Level Code Generation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
VulAgent: Hypothesis-Validation based Multi-Agent Vulnerability Detection
by: Wang, Ziliang, et al.
Published: (2025)
by: Wang, Ziliang, et al.
Published: (2025)
FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning
by: Ding, Haoran, et al.
Published: (2026)
by: Ding, Haoran, et al.
Published: (2026)
Environment-in-the-Loop: Rethinking Code Migration with LLM-based Agents
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Empirical Research on Utilizing LLM-based Agents for Automated Bug Fixing via LangGraph
by: Wang, Jialin, et al.
Published: (2025)
by: Wang, Jialin, et al.
Published: (2025)
Towards Engineering Multi-Agent LLMs: A Protocol-Driven Approach
by: Mao, Zhenyu, et al.
Published: (2025)
by: Mao, Zhenyu, et al.
Published: (2025)
LLM-Agents Driven Automated Simulation Testing and Analysis of small Uncrewed Aerial Systems
by: Duvvuru, Venkata Sai Aswath, et al.
Published: (2025)
by: Duvvuru, Venkata Sai Aswath, et al.
Published: (2025)
MASPrism: Lightweight Failure Attribution for Multi-Agent Systems Using Prefill-Stage Signals
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
GeoJSON Agents:A Multi-Agent LLM Architecture for Geospatial Analysis-Function Calling vs Code Generation
by: Luo, Qianqian, et al.
Published: (2025)
by: Luo, Qianqian, et al.
Published: (2025)
From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems
by: Wang, Yawen, et al.
Published: (2026)
by: Wang, Yawen, et al.
Published: (2026)
Similar Items
-
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
by: Zhang, Lingzhe, et al.
Published: (2026) -
LogDB: Multivariate Log-based Failure Diagnosis for Distributed Databases (Extended from MultiLog)
by: Zhang, Lingzhe, et al.
Published: (2025) -
ThinkFL: Self-Refining Failure Localization for Microservice Systems via Reinforcement Fine-Tuning
by: Zhang, Lingzhe, et al.
Published: (2025) -
Efficient Failure Management for Multi-Agent Systems with Reasoning Trace Representation
by: Zhang, Lingzhe, et al.
Published: (2026) -
Adaptive Root Cause Localization for Microservice Systems with Multi-Agent Recursion-of-Thought
by: Zhang, Lingzhe, et al.
Published: (2025)