Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level Rectification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Yaoyang, Zheng, Zhi, Zhao, Ziwei, Xu, Tong, Jielun, Zhao, Xue, Wenjun, Chen, Yong, Chen, Enhong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917538523250688
author Luo, Yaoyang
Zheng, Zhi
Zhao, Ziwei
Xu, Tong
Jielun, Zhao
Xue, Wenjun
Chen, Yong
Chen, Enhong
author_facet Luo, Yaoyang
Zheng, Zhi
Zhao, Ziwei
Xu, Tong
Jielun, Zhao
Xue, Wenjun
Chen, Yong
Chen, Enhong
contents Recent years have witnessed the rapid development of Large Language Model-based Multi-Agent Systems (MAS), which excel at collaborative decision-making and complex problem-solving. However, malicious agents in MAS may inject misinformation to mislead other agents and disrupt system performance, giving rise to a new research direction that focuses on attack mechanisms and defense strategies in MAS. Prior studies largely assume malicious agents act independently and investigate the corresponding defense strategies. However, we argue that malicious agents may exhibit collaborative behaviors, enabling more effective attacks through internal information exchange. In this paper, we propose an adaptive cooperative attack framework, where malicious agents autonomously coordinate and dynamically adjust their attack strategies through multi-round interactions. Furthermore, we introduce Sentence-Level Trustworthiness Analysis and Rectification (STAR), a defense framework that identifies and rectifies misleading information at the sentence level within agent communications. Our experiments show that cooperative attacks lead to a significantly larger degradation in task success rate than independent attacks, resulting in a relative drop of 5.34\%. Meanwhile, STAR effectively mitigates both cooperative and independent threats and improves task success rate by an average of 36.76\%. The code is available at https://github.com/smoooom/STAR.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28104
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level Rectification
Luo, Yaoyang
Zheng, Zhi
Zhao, Ziwei
Xu, Tong
Jielun, Zhao
Xue, Wenjun
Chen, Yong
Chen, Enhong
Artificial Intelligence
Recent years have witnessed the rapid development of Large Language Model-based Multi-Agent Systems (MAS), which excel at collaborative decision-making and complex problem-solving. However, malicious agents in MAS may inject misinformation to mislead other agents and disrupt system performance, giving rise to a new research direction that focuses on attack mechanisms and defense strategies in MAS. Prior studies largely assume malicious agents act independently and investigate the corresponding defense strategies. However, we argue that malicious agents may exhibit collaborative behaviors, enabling more effective attacks through internal information exchange. In this paper, we propose an adaptive cooperative attack framework, where malicious agents autonomously coordinate and dynamically adjust their attack strategies through multi-round interactions. Furthermore, we introduce Sentence-Level Trustworthiness Analysis and Rectification (STAR), a defense framework that identifies and rectifies misleading information at the sentence level within agent communications. Our experiments show that cooperative attacks lead to a significantly larger degradation in task success rate than independent attacks, resulting in a relative drop of 5.34\%. Meanwhile, STAR effectively mitigates both cooperative and independent threats and improves task success rate by an average of 36.76\%. The code is available at https://github.com/smoooom/STAR.
title Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level Rectification
topic Artificial Intelligence
url https://arxiv.org/abs/2605.28104