MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety
Fuente:
arXiv
Salvato in:
| Autori principali: | Wen, Xiaoyu, He, Zhida, Qi, Han, Wan, Ziyu, Ma, Zhongtian, Wen, Ying, Zheng, Tianhang, Xu, Xingcheng, Lu, Chaochao, Zhang, Qiaosheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking
di: He, Zhida, et al.
Pubblicazione: (2026)
di: He, Zhida, et al.
Pubblicazione: (2026)
Multi-Defender Single-Attacker Perimeter Defense Game on a Cylinder: Special Case in which the Attacker Starts at the Boundary
di: Otte, Michael, et al.
Pubblicazione: (2026)
di: Otte, Michael, et al.
Pubblicazione: (2026)
Language Games as the Pathway to Artificial Superhuman Intelligence
di: Wen, Ying, et al.
Pubblicazione: (2025)
di: Wen, Ying, et al.
Pubblicazione: (2025)
IDCAIS: Inter-Defender Collision-Aware Interception Strategy against Multiple Attackers
di: Chipade, Vishnu S., et al.
Pubblicazione: (2021)
di: Chipade, Vishnu S., et al.
Pubblicazione: (2021)
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
di: Jin, Chang, et al.
Pubblicazione: (2026)
di: Jin, Chang, et al.
Pubblicazione: (2026)
Defending a City from Multi-Drone Attacks: A Sequential Stackelberg Security Games Approach
di: Mutzari, Dolev, et al.
Pubblicazione: (2025)
di: Mutzari, Dolev, et al.
Pubblicazione: (2025)
CoTDeceptor:Adversarial Code Obfuscation Against CoT-Enhanced LLM Code Agents
di: Li, Haoyang, et al.
Pubblicazione: (2025)
di: Li, Haoyang, et al.
Pubblicazione: (2025)
Dynamic Adversarial Resource Allocation: the dDAB Game
di: Guan, Yue, et al.
Pubblicazione: (2023)
di: Guan, Yue, et al.
Pubblicazione: (2023)
ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
di: Wan, Ziyu, et al.
Pubblicazione: (2025)
di: Wan, Ziyu, et al.
Pubblicazione: (2025)
Constitutional Arms Races in the Public Goods Game: Co-Evolving LLM Constitutions Under Cooperation-Defection Pressure
di: Kumar, Ujwal, et al.
Pubblicazione: (2026)
di: Kumar, Ujwal, et al.
Pubblicazione: (2026)
Guarding a Target Area from a Heterogeneous Group of Cooperative Attackers
di: Lee, Yoonjae, et al.
Pubblicazione: (2024)
di: Lee, Yoonjae, et al.
Pubblicazione: (2024)
Aegis:An Advanced LLM-Based Multi-Agent for Intelligent Functional Safety Engineering
di: Shi, Lu, et al.
Pubblicazione: (2024)
di: Shi, Lu, et al.
Pubblicazione: (2024)
Hierarchical Decision-Making in Population Games
di: Chen, Yu-Wen, et al.
Pubblicazione: (2025)
di: Chen, Yu-Wen, et al.
Pubblicazione: (2025)
Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment
di: Li, Jiajia, et al.
Pubblicazione: (2026)
di: Li, Jiajia, et al.
Pubblicazione: (2026)
STLGame: Signal Temporal Logic Games in Adversarial Multi-Agent Systems
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
Leveraging Team Correlation for Approximating Equilibrium in Two-Team Zero-Sum Games
di: Liu, Naming, et al.
Pubblicazione: (2024)
di: Liu, Naming, et al.
Pubblicazione: (2024)
MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning
di: Yu, Haohan, et al.
Pubblicazione: (2026)
di: Yu, Haohan, et al.
Pubblicazione: (2026)
Robust Multi-agent Communication Based on Decentralization-Oriented Adversarial Training
di: Ma, Xuyan, et al.
Pubblicazione: (2025)
di: Ma, Xuyan, et al.
Pubblicazione: (2025)
MermaidFlow: Redefining Agentic Workflow Generation via Safety-Constrained Evolutionary Programming
di: Zheng, Chengqi, et al.
Pubblicazione: (2025)
di: Zheng, Chengqi, et al.
Pubblicazione: (2025)
Multi-LLM Systems Exhibit Robust Semantic Collapse
di: Kong, Weiyi, et al.
Pubblicazione: (2026)
di: Kong, Weiyi, et al.
Pubblicazione: (2026)
Conformity Dynamics in LLM Multi-Agent Systems: The Roles of Topology and Self-Social Weighting
di: Han, Chen, et al.
Pubblicazione: (2026)
di: Han, Chen, et al.
Pubblicazione: (2026)
A Spatial Calibration Method for Robust Cooperative Perception
di: Song, Zhiying, et al.
Pubblicazione: (2023)
di: Song, Zhiying, et al.
Pubblicazione: (2023)
Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems
di: Hao, Zhezheng, et al.
Pubblicazione: (2026)
di: Hao, Zhezheng, et al.
Pubblicazione: (2026)
Offline Fictitious Self-Play for Competitive Games
di: Chen, Jingxiao, et al.
Pubblicazione: (2024)
di: Chen, Jingxiao, et al.
Pubblicazione: (2024)
Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement Learning
di: Lee, Sunwoo, et al.
Pubblicazione: (2026)
di: Lee, Sunwoo, et al.
Pubblicazione: (2026)
Safe Multi-Agent Behavior Must Be Maintained, Not Merely Asserted: Constraint Drift in LLM-Based Multi-Agent Systems
di: Li, Tianxiao, et al.
Pubblicazione: (2026)
di: Li, Tianxiao, et al.
Pubblicazione: (2026)
Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions
di: Bisconti, Piercosma, et al.
Pubblicazione: (2025)
di: Bisconti, Piercosma, et al.
Pubblicazione: (2025)
LLM-ALSO: LLM-Driven Adaptive Learning-Signal Optimization for Multi-Agent Reinforcement Learning
di: Wu, Xiaoguang, et al.
Pubblicazione: (2026)
di: Wu, Xiaoguang, et al.
Pubblicazione: (2026)
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
di: Kavathekar, Ishan, et al.
Pubblicazione: (2025)
di: Kavathekar, Ishan, et al.
Pubblicazione: (2025)
Robust Multi-Agent Decision-Making in Finite-Population Games
di: Park, Shinkyu, et al.
Pubblicazione: (2025)
di: Park, Shinkyu, et al.
Pubblicazione: (2025)
Game-Theoretic Lens on LLM-based Multi-Agent Systems
di: Hao, Jianing, et al.
Pubblicazione: (2026)
di: Hao, Jianing, et al.
Pubblicazione: (2026)
Agent Exchange: Shaping the Future of AI Agent Economics
di: Yang, Yingxuan, et al.
Pubblicazione: (2025)
di: Yang, Yingxuan, et al.
Pubblicazione: (2025)
Diffusion-Reinforcement Learning Hierarchical Motion Planning in Multi-agent Adversarial Games
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
Adaptive Orchestration: Scalable Self-Evolving Multi-Agent Systems
di: Sampath, Sathish, et al.
Pubblicazione: (2026)
di: Sampath, Sathish, et al.
Pubblicazione: (2026)
Self-Evolving Multi-Agent Systems via Decentralized Memory
di: Hao, Guangya, et al.
Pubblicazione: (2026)
di: Hao, Guangya, et al.
Pubblicazione: (2026)
Open-Ended Video Game Glitch Detection with Agentic Reasoning and Temporal Grounding
di: Zheng, Muyang, et al.
Pubblicazione: (2026)
di: Zheng, Muyang, et al.
Pubblicazione: (2026)
Provably Efficient Information-Directed Sampling Algorithms for Multi-Agent Reinforcement Learning
di: Zhang, Qiaosheng, et al.
Pubblicazione: (2024)
di: Zhang, Qiaosheng, et al.
Pubblicazione: (2024)
DialogGuard: Multi-Agent Psychosocial Safety Evaluation of Sensitive LLM Responses
di: Luo, Han, et al.
Pubblicazione: (2025)
di: Luo, Han, et al.
Pubblicazione: (2025)
TrafficGamer: Reliable and Flexible Traffic Simulation for Safety-Critical Scenarios with Game-Theoretic Oracles
di: Qiao, Guanren, et al.
Pubblicazione: (2024)
di: Qiao, Guanren, et al.
Pubblicazione: (2024)
NetSafe: Exploring the Topological Safety of Multi-agent Networks
di: Yu, Miao, et al.
Pubblicazione: (2024)
di: Yu, Miao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking
di: He, Zhida, et al.
Pubblicazione: (2026) -
Multi-Defender Single-Attacker Perimeter Defense Game on a Cylinder: Special Case in which the Attacker Starts at the Boundary
di: Otte, Michael, et al.
Pubblicazione: (2026) -
Language Games as the Pathway to Artificial Superhuman Intelligence
di: Wen, Ying, et al.
Pubblicazione: (2025) -
IDCAIS: Inter-Defender Collision-Aware Interception Strategy against Multiple Attackers
di: Chipade, Vishnu S., et al.
Pubblicazione: (2021) -
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
di: Jin, Chang, et al.
Pubblicazione: (2026)