BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miao, Rui, Liu, Yixin, Wang, Yili, Shen, Xu, Tan, Yue, Dai, Yiwei, Pan, Shirui, Wang, Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914507121491968
author Miao, Rui
Liu, Yixin
Wang, Yili
Shen, Xu
Tan, Yue
Dai, Yiwei
Pan, Shirui
Wang, Xin
author_facet Miao, Rui
Liu, Yixin
Wang, Yili
Shen, Xu
Tan, Yue
Dai, Yiwei
Pan, Shirui
Wang, Xin
contents The security of LLM-based multi-agent systems (MAS) is critically threatened by propagation vulnerability, where malicious agents can distort collective decision-making through inter-agent message interactions. While existing supervised defense methods demonstrate promising performance, they may be impractical in real-world scenarios due to their heavy reliance on labeled malicious agents to train a supervised malicious detection model. To enable practical and generalizable MAS defenses, in this paper, we propose BlindGuard, an unsupervised defense method that learns without requiring any attack-specific labels or prior knowledge of malicious behaviors. To this end, we establish a hierarchical agent encoder to capture individual, neighborhood, and global interaction patterns of each agent, providing a comprehensive understanding for malicious agent detection. Meanwhile, we design a corruption-guided detector that consists of directional noise injection and contrastive learning, allowing effective detection model training solely on normal agent behaviors. Extensive experiments show that BlindGuard effectively detects diverse attack types (i.e., prompt injection, memory poisoning, and tool attack) across MAS with various communication patterns while maintaining superior generalizability compared to supervised baselines. The code is available at: https://github.com/MR9812/BlindGuard.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08127
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
Miao, Rui
Liu, Yixin
Wang, Yili
Shen, Xu
Tan, Yue
Dai, Yiwei
Pan, Shirui
Wang, Xin
Artificial Intelligence
The security of LLM-based multi-agent systems (MAS) is critically threatened by propagation vulnerability, where malicious agents can distort collective decision-making through inter-agent message interactions. While existing supervised defense methods demonstrate promising performance, they may be impractical in real-world scenarios due to their heavy reliance on labeled malicious agents to train a supervised malicious detection model. To enable practical and generalizable MAS defenses, in this paper, we propose BlindGuard, an unsupervised defense method that learns without requiring any attack-specific labels or prior knowledge of malicious behaviors. To this end, we establish a hierarchical agent encoder to capture individual, neighborhood, and global interaction patterns of each agent, providing a comprehensive understanding for malicious agent detection. Meanwhile, we design a corruption-guided detector that consists of directional noise injection and contrastive learning, allowing effective detection model training solely on normal agent behaviors. Extensive experiments show that BlindGuard effectively detects diverse attack types (i.e., prompt injection, memory poisoning, and tool attack) across MAS with various communication patterns while maintaining superior generalizability compared to supervised baselines. The code is available at: https://github.com/MR9812/BlindGuard.
title BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
topic Artificial Intelligence
url https://arxiv.org/abs/2508.08127