TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Kai, Zeng, Biaojie, Wei, Zeming, Jin, Chang, Zhou, Hefeng, Li, Xiangtian, Yang, Chao, Qu, Jingjing, Xu, Xingcheng, Hu, Xia
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918391443357696
author Wang, Kai
Zeng, Biaojie
Wei, Zeming
Jin, Chang
Zhou, Hefeng
Li, Xiangtian
Yang, Chao
Qu, Jingjing
Xu, Xingcheng
Hu, Xia
author_facet Wang, Kai
Zeng, Biaojie
Wei, Zeming
Jin, Chang
Zhou, Hefeng
Li, Xiangtian
Yang, Chao
Qu, Jingjing
Xu, Xingcheng
Hu, Xia
contents With the rapid development of LLM-based multi-agent systems (MAS), their significant safety and security concerns have emerged, which introduce novel risks going beyond single agents or LLMs. Despite attempts to address these issues, the existing literature lacks a cohesive safeguarding system specialized for MAS risks. In this work, we introduce TrinityGuard, a comprehensive safety evaluation and monitoring framework for LLM-based MAS, grounded in the OWASP standards. Specifically, TrinityGuard encompasses a three-tier fine-grained risk taxonomy that identifies 20 risk types, covering single-agent vulnerabilities, inter-agent communication threats, and system-level emergent hazards. Designed for scalability across various MAS structures and platforms, TrinityGuard is organized in a trinity manner, involving an MAS abstraction layer that can be adapted to any MAS structures, an evaluation layer containing risk-specific test modules, alongside runtime monitor agents coordinated by a unified LLM Judge Factory. During Evaluation, TrinityGuard executes curated attack probes to generate detailed vulnerability reports for each risk type, where monitor agents analyze structured execution traces and issue real-time alerts, enabling both pre-development evaluation and runtime monitoring. We further formalize these safety metrics and present detailed case studies across various representative MAS examples, showcasing the versatility and reliability of TrinityGuard. Overall, TrinityGuard acts as a comprehensive framework for evaluating and monitoring various risks in MAS, paving the way for further research into their safety and security.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15408
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems
Wang, Kai
Zeng, Biaojie
Wei, Zeming
Jin, Chang
Zhou, Hefeng
Li, Xiangtian
Yang, Chao
Qu, Jingjing
Xu, Xingcheng
Hu, Xia
Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
Multiagent Systems
With the rapid development of LLM-based multi-agent systems (MAS), their significant safety and security concerns have emerged, which introduce novel risks going beyond single agents or LLMs. Despite attempts to address these issues, the existing literature lacks a cohesive safeguarding system specialized for MAS risks. In this work, we introduce TrinityGuard, a comprehensive safety evaluation and monitoring framework for LLM-based MAS, grounded in the OWASP standards. Specifically, TrinityGuard encompasses a three-tier fine-grained risk taxonomy that identifies 20 risk types, covering single-agent vulnerabilities, inter-agent communication threats, and system-level emergent hazards. Designed for scalability across various MAS structures and platforms, TrinityGuard is organized in a trinity manner, involving an MAS abstraction layer that can be adapted to any MAS structures, an evaluation layer containing risk-specific test modules, alongside runtime monitor agents coordinated by a unified LLM Judge Factory. During Evaluation, TrinityGuard executes curated attack probes to generate detailed vulnerability reports for each risk type, where monitor agents analyze structured execution traces and issue real-time alerts, enabling both pre-development evaluation and runtime monitoring. We further formalize these safety metrics and present detailed case studies across various representative MAS examples, showcasing the versatility and reliability of TrinityGuard. Overall, TrinityGuard acts as a comprehensive framework for evaluating and monitoring various risks in MAS, paving the way for further research into their safety and security.
title TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2603.15408