When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Lingxi, Zheng, Guangtao, Chen, Hanjie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explainable Autonomous Cyber Defense using Adversarial Multi-Agent Reinforcement Learning
by: Zhang, Yiyao, et al.
Published: (2026)
by: Zhang, Yiyao, et al.
Published: (2026)
Multi-Agent Framework for Threat Mitigation and Resilience in AI-Based Systems
by: Foundjem, Armstrong, et al.
Published: (2025)
by: Foundjem, Armstrong, et al.
Published: (2025)
Information-Theoretic Privacy Control for Sequential Multi-Agent LLM Systems
by: Asif, Sadia, et al.
Published: (2026)
by: Asif, Sadia, et al.
Published: (2026)
Memory Poisoning Attack and Defense on Memory Based LLM-Agents
by: Sunil, Balachandra Devarangadi, et al.
Published: (2026)
by: Sunil, Balachandra Devarangadi, et al.
Published: (2026)
Hierarchical Multi-agent Reinforcement Learning for Cyber Network Defense
by: Singh, Aditya Vikram, et al.
Published: (2024)
by: Singh, Aditya Vikram, et al.
Published: (2024)
Many-to-One Adversarial Consensus: Exposing Multi-Agent Collusion Risks in AI-Based Healthcare
by: Bashir, Adeela, et al.
Published: (2025)
by: Bashir, Adeela, et al.
Published: (2025)
G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent Systems
by: Wang, Shilong, et al.
Published: (2025)
by: Wang, Shilong, et al.
Published: (2025)
Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems
by: Wang, Chenxi, et al.
Published: (2026)
by: Wang, Chenxi, et al.
Published: (2026)
Quantitative Resilience Modeling for Autonomous Cyber Defense
by: Cadet, Xavier, et al.
Published: (2025)
by: Cadet, Xavier, et al.
Published: (2025)
CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution
by: Shao, Minghao, et al.
Published: (2025)
by: Shao, Minghao, et al.
Published: (2025)
Architecture Matters for Multi-Agent Security
by: Hagag, Ben, et al.
Published: (2026)
by: Hagag, Ben, et al.
Published: (2026)
CEE: An Inference-Time Jailbreak Defense for Embodied Intelligence via Subspace Concept Rotation
by: Yang, Jirui, et al.
Published: (2025)
by: Yang, Jirui, et al.
Published: (2025)
Multi-Agent Reinforcement Learning for Maritime Operational Technology Cyber Security
by: Wilson, Alec, et al.
Published: (2024)
by: Wilson, Alec, et al.
Published: (2024)
Learning to Communicate in Multi-Agent Reinforcement Learning for Autonomous Cyber Defence
by: Contractor, Faizan, et al.
Published: (2025)
by: Contractor, Faizan, et al.
Published: (2025)
Scalable Multi-Agent Reinforcement Learning for Residential Load Scheduling under Data Governance
by: Qin, Zhaoming, et al.
Published: (2021)
by: Qin, Zhaoming, et al.
Published: (2021)
CuDA2: An approach for Incorporating Traitor Agents into Cooperative Multi-Agent Systems
by: Chen, Zhen, et al.
Published: (2024)
by: Chen, Zhen, et al.
Published: (2024)
MAGIQ: A Post-Quantum Multi-Agentic AI Governance System with Provable Security
by: Avizheh, Sepideh, et al.
Published: (2026)
by: Avizheh, Sepideh, et al.
Published: (2026)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
by: Rahman, Salman, et al.
Published: (2025)
by: Rahman, Salman, et al.
Published: (2025)
Beyond Single-Agent Alignment: Preventing Context-Fragmented Violations in Multi-Agent Systems
by: Wu, Jie, et al.
Published: (2026)
by: Wu, Jie, et al.
Published: (2026)
Breaking the Secret: Economic Interventions for Combating Collusion in Embodied Multi-Agent Systems
by: Liu, Qi, et al.
Published: (2026)
by: Liu, Qi, et al.
Published: (2026)
Hierarchical Adversarially-Resilient Multi-Agent Reinforcement Learning for Cyber-Physical Systems Security
by: Alqithami, Saad
Published: (2025)
by: Alqithami, Saad
Published: (2025)
Towards Transparent and Incentive-Compatible Collaboration in Decentralized LLM Multi-Agent Systems: A Blockchain-Driven Approach
by: Qi, Minfeng, et al.
Published: (2025)
by: Qi, Minfeng, et al.
Published: (2025)
Learning Communication Between Heterogeneous Agents in Multi-Agent Reinforcement Learning for Autonomous Cyber Defence
by: Popa, Alex, et al.
Published: (2026)
by: Popa, Alex, et al.
Published: (2026)
Locally Differentially Private Distributed Online Learning with Guaranteed Optimality
by: Chen, Ziqin, et al.
Published: (2023)
by: Chen, Ziqin, et al.
Published: (2023)
Wolfpack Adversarial Attack for Robust Multi-Agent Reinforcement Learning
by: Lee, Sunwoo, et al.
Published: (2025)
by: Lee, Sunwoo, et al.
Published: (2025)
Optimal Cost Constrained Adversarial Attacks For Multiple Agent Systems
by: Lu, Ziqing, et al.
Published: (2023)
by: Lu, Ziqing, et al.
Published: (2023)
Decentralized Multi-Agent System with Trust-Aware Communication
by: Ding, Yepeng, et al.
Published: (2025)
by: Ding, Yepeng, et al.
Published: (2025)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
by: Jin, Chang, et al.
Published: (2026)
by: Jin, Chang, et al.
Published: (2026)
Secure Forgetting: A Framework for Privacy-Driven Unlearning in Large Language Model (LLM)-Based Agents
by: Ye, Dayong, et al.
Published: (2026)
by: Ye, Dayong, et al.
Published: (2026)
GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems
by: Mateo-Torrejón, Pablo, et al.
Published: (2026)
by: Mateo-Torrejón, Pablo, et al.
Published: (2026)
Multi-Agent Actor-Critics in Autonomous Cyber Defense
by: Wang, Mingjun, et al.
Published: (2024)
by: Wang, Mingjun, et al.
Published: (2024)
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
Attacks and Mitigations for Distributed Governance of Agentic AI under Byzantine Adversaries
by: Laws, Matthew D., et al.
Published: (2026)
by: Laws, Matthew D., et al.
Published: (2026)
Optimizing Day-Ahead Energy Trading with Proximal Policy Optimization and Blockchain
by: Verma, Navneet, et al.
Published: (2025)
by: Verma, Navneet, et al.
Published: (2025)
Enhancing the Robustness of QMIX against State-adversarial Attacks
by: Guo, Weiran, et al.
Published: (2023)
by: Guo, Weiran, et al.
Published: (2023)
ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
by: Zhang, Yihao, et al.
Published: (2026)
by: Zhang, Yihao, et al.
Published: (2026)
TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
The Best-Laid SCHEMEs: Coordinated Sabotage and Monitoring in Multi-Agent Systems
by: Radev, Nikolay, et al.
Published: (2026)
by: Radev, Nikolay, et al.
Published: (2026)
Web Fraud Attacks Against LLM-Driven Multi-Agent Systems
by: Kong, Dezhang, et al.
Published: (2025)
by: Kong, Dezhang, et al.
Published: (2025)
Defending Against Diverse Attacks in Federated Learning Through Consensus-Based Bi-Level Optimization
by: Trillos, Nicolás García, et al.
Published: (2024)
by: Trillos, Nicolás García, et al.
Published: (2024)
Similar Items
-
Explainable Autonomous Cyber Defense using Adversarial Multi-Agent Reinforcement Learning
by: Zhang, Yiyao, et al.
Published: (2026) -
Multi-Agent Framework for Threat Mitigation and Resilience in AI-Based Systems
by: Foundjem, Armstrong, et al.
Published: (2025) -
Information-Theoretic Privacy Control for Sequential Multi-Agent LLM Systems
by: Asif, Sadia, et al.
Published: (2026) -
Memory Poisoning Attack and Defense on Memory Based LLM-Agents
by: Sunil, Balachandra Devarangadi, et al.
Published: (2026) -
Hierarchical Multi-agent Reinforcement Learning for Cyber Network Defense
by: Singh, Aditya Vikram, et al.
Published: (2024)