Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level Rectification
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Yaoyang, Zheng, Zhi, Zhao, Ziwei, Xu, Tong, Jielun, Zhao, Xue, Wenjun, Chen, Yong, Chen, Enhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DynaDebate: Breaking Homogeneity in Multi-Agent Debate with Dynamic Path Generation
by: Li, Zhenghao, et al.
Published: (2026)
by: Li, Zhenghao, et al.
Published: (2026)
VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
by: Lin, Junda, et al.
Published: (2026)
by: Lin, Junda, et al.
Published: (2026)
Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data
by: Dong, Fengxian, et al.
Published: (2026)
by: Dong, Fengxian, et al.
Published: (2026)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023)
by: Cao, Bochuan, et al.
Published: (2023)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
No Free Lunch for Defending Against Prefilling Attack by In-Context Learning
by: Xue, Zhiyu, et al.
Published: (2024)
by: Xue, Zhiyu, et al.
Published: (2024)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
by: Zhong, Peter Yong, et al.
Published: (2025)
by: Zhong, Peter Yong, et al.
Published: (2025)
Web Fraud Attacks Against LLM-Driven Multi-Agent Systems
by: Kong, Dezhang, et al.
Published: (2025)
by: Kong, Dezhang, et al.
Published: (2025)
PeerGuard: Defending Multi-Agent Systems Against Backdoor Attacks Through Mutual Reasoning
by: Fan, Falong, et al.
Published: (2025)
by: Fan, Falong, et al.
Published: (2025)
CuDA2: An approach for Incorporating Traitor Agents into Cooperative Multi-Agent Systems
by: Chen, Zhen, et al.
Published: (2024)
by: Chen, Zhen, et al.
Published: (2024)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
DynLLM: When Large Language Models Meet Dynamic Graph Recommendation
by: Zhao, Ziwei, et al.
Published: (2024)
by: Zhao, Ziwei, et al.
Published: (2024)
Defending Against Network Attacks for Secure AI Agent Migration in Vehicular Metaverses
by: Wen, Xinru, et al.
Published: (2024)
by: Wen, Xinru, et al.
Published: (2024)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
by: Liu, Tiantian, et al.
Published: (2024)
by: Liu, Tiantian, et al.
Published: (2024)
Defending against Backdoor Attack on Deep Neural Networks
by: Cheng, Hao, et al.
Published: (2020)
by: Cheng, Hao, et al.
Published: (2020)
GroupGuard: A Framework for Modeling and Defending Collusive Attacks in Multi-Agent Systems
by: Tao, Yiling, et al.
Published: (2026)
by: Tao, Yiling, et al.
Published: (2026)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
by: Jia, Feiran, et al.
Published: (2024)
by: Jia, Feiran, et al.
Published: (2024)
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
by: Robey, Alexander, et al.
Published: (2023)
by: Robey, Alexander, et al.
Published: (2023)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
by: An, Li, et al.
Published: (2025)
by: An, Li, et al.
Published: (2025)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
by: Zhang, Dongsen, et al.
Published: (2025)
by: Zhang, Dongsen, et al.
Published: (2025)
Defending Against Social Engineering Attacks in the Age of LLMs
by: Ai, Lin, et al.
Published: (2024)
by: Ai, Lin, et al.
Published: (2024)
To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack
by: Zhuo, Terry Yue, et al.
Published: (2026)
by: Zhuo, Terry Yue, et al.
Published: (2026)
Defending Against Data Reconstruction Attacks in Federated Learning: An Information Theory Approach
by: Tan, Qi, et al.
Published: (2024)
by: Tan, Qi, et al.
Published: (2024)
Direct interpolative construction of the discrete Fourier transform as a matrix product operator
by: Chen, Jielun, et al.
Published: (2024)
by: Chen, Jielun, et al.
Published: (2024)
Fight Fire with Fire: Defending Against Malicious RL Fine-Tuning via Reward Neutralization
by: Cao, Wenjun
Published: (2025)
by: Cao, Wenjun
Published: (2025)
Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
Defending Against Physical Adversarial Patch Attacks on Infrared Human Detection
by: Strack, Lukas, et al.
Published: (2023)
by: Strack, Lukas, et al.
Published: (2023)
Attacking Cooperative Multi-Agent Reinforcement Learning by Adversarial Minority Influence
by: Li, Simin, et al.
Published: (2023)
by: Li, Simin, et al.
Published: (2023)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
Prompt-Unknown Promotion Attacks against LLM-based Sequential Recommender Systems
by: Zhao, Yuchuan, et al.
Published: (2026)
by: Zhao, Yuchuan, et al.
Published: (2026)
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
by: Ying, Zonghao, et al.
Published: (2026)
by: Ying, Zonghao, et al.
Published: (2026)
DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Defending Against Diverse Attacks in Federated Learning Through Consensus-Based Bi-Level Optimization
by: Trillos, Nicolás García, et al.
Published: (2024)
by: Trillos, Nicolás García, et al.
Published: (2024)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
by: Xin, Yuan, et al.
Published: (2026)
by: Xin, Yuan, et al.
Published: (2026)
Constrained Black-Box Attacks Against Cooperative Multi-Agent Reinforcement Learning
by: Andam, Amine, et al.
Published: (2025)
by: Andam, Amine, et al.
Published: (2025)
Filter, Obstruct and Dilute: Defending Against Backdoor Attacks on Semi-Supervised Learning
by: Wang, Xinrui, et al.
Published: (2025)
by: Wang, Xinrui, et al.
Published: (2025)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
by: Hines, Keegan, et al.
Published: (2024)
by: Hines, Keegan, et al.
Published: (2024)
Defending Against Frequency-Based Attacks with Diffusion Models
by: Amerehi, Fatemeh, et al.
Published: (2025)
by: Amerehi, Fatemeh, et al.
Published: (2025)
Similar Items
-
DynaDebate: Breaking Homogeneity in Multi-Agent Debate with Dynamic Path Generation
by: Li, Zhenghao, et al.
Published: (2026) -
VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
by: Lin, Junda, et al.
Published: (2026) -
Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data
by: Dong, Fengxian, et al.
Published: (2026) -
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023) -
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)