The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
Fuente:
arXiv
Salvato in:
| Autori principali: | La Malfa, Gabriele, La Malfa, Emanuele, Cohen, Saar, Zhang, Jie M., Luck, Michael, Wooldridge, Michael, Black, Elizabeth |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tacit Coordination of Large Language Models
di: Aharon, Ido, et al.
Pubblicazione: (2026)
di: Aharon, Ido, et al.
Pubblicazione: (2026)
Large Language Models Miss the Multi-Agent Mark
di: La Malfa, Emanuele, et al.
Pubblicazione: (2025)
di: La Malfa, Emanuele, et al.
Pubblicazione: (2025)
Using Protected Attributes to Consider Fairness in Multi-Agent Systems
di: La Malfa, Gabriele, et al.
Pubblicazione: (2024)
di: La Malfa, Gabriele, et al.
Pubblicazione: (2024)
Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI
di: Huang, Xuanqiang Angelo, et al.
Pubblicazione: (2026)
di: Huang, Xuanqiang Angelo, et al.
Pubblicazione: (2026)
Deep Neural Networks via Complex Network Theory: a Perspective
di: La Malfa, Emanuele, et al.
Pubblicazione: (2024)
di: La Malfa, Emanuele, et al.
Pubblicazione: (2024)
Fairness Aware Reinforcement Learning via Proximal Policy Optimization
di: La Malfa, Gabriele, et al.
Pubblicazione: (2025)
di: La Malfa, Gabriele, et al.
Pubblicazione: (2025)
Fixed Point Explainability
di: La Malfa, Emanuele, et al.
Pubblicazione: (2025)
di: La Malfa, Emanuele, et al.
Pubblicazione: (2025)
Offline Learning of Nash Stable Coalition Structures with Possibly Overlapping Coalitions
di: Cohen, Saar
Pubblicazione: (2026)
di: Cohen, Saar
Pubblicazione: (2026)
End-to-end PDDL Planning with Hardcoded and Dynamic Agents
di: La Malfa, Emanuele, et al.
Pubblicazione: (2025)
di: La Malfa, Emanuele, et al.
Pubblicazione: (2025)
Characterising Interventions in Causal Games
di: Mishra, Manuj, et al.
Pubblicazione: (2024)
di: Mishra, Manuj, et al.
Pubblicazione: (2024)
Multi-agent coordination via communication partitions
di: Lee, Wei-Chen, et al.
Pubblicazione: (2025)
di: Lee, Wei-Chen, et al.
Pubblicazione: (2025)
Out-of-Context Reasoning in Large Language Models
di: Shaki, Jonathan, et al.
Pubblicazione: (2025)
di: Shaki, Jonathan, et al.
Pubblicazione: (2025)
Language Self-Play For Data-Free Training
di: Kuba, Jakub Grudzien, et al.
Pubblicazione: (2025)
di: Kuba, Jakub Grudzien, et al.
Pubblicazione: (2025)
Quantifying the Self-Interest Level of Markov Social Dilemmas
di: Willis, Richard, et al.
Pubblicazione: (2025)
di: Willis, Richard, et al.
Pubblicazione: (2025)
A Notion of Complexity for Theory of Mind via Discrete World Models
di: Huang, X. Angelo, et al.
Pubblicazione: (2024)
di: Huang, X. Angelo, et al.
Pubblicazione: (2024)
GAE Falls Short in Imperfect-Information Self-Play Reinforcement Learning
di: Fan, Zhiyuan, et al.
Pubblicazione: (2026)
di: Fan, Zhiyuan, et al.
Pubblicazione: (2026)
Game Networks
di: La Mura, Pierfrancesco
Pubblicazione: (2013)
di: La Mura, Pierfrancesco
Pubblicazione: (2013)
Meta-Learning in Self-Play Regret Minimization
di: Sychrovský, David, et al.
Pubblicazione: (2025)
di: Sychrovský, David, et al.
Pubblicazione: (2025)
Fetch.ai: An Architecture for Modern Multi-Agent Systems
di: Wooldridge, Michael J., et al.
Pubblicazione: (2025)
di: Wooldridge, Michael J., et al.
Pubblicazione: (2025)
Offline Fictitious Self-Play for Competitive Games
di: Chen, Jingxiao, et al.
Pubblicazione: (2024)
di: Chen, Jingxiao, et al.
Pubblicazione: (2024)
Expected Utility Networks
di: La Mura, Pierfrancesco, et al.
Pubblicazione: (2013)
di: La Mura, Pierfrancesco, et al.
Pubblicazione: (2013)
Fixed-budget and Multiple-issue Quadratic Voting
di: Georgescu, Laura, et al.
Pubblicazione: (2024)
di: Georgescu, Laura, et al.
Pubblicazione: (2024)
Pure Exploration via Frank-Wolfe Self-Play
di: Liu, Xinyu, et al.
Pubblicazione: (2025)
di: Liu, Xinyu, et al.
Pubblicazione: (2025)
Adaptive Honeypot Allocation in Multi-Attacker Networks via Bayesian Stackelberg Games
di: Park, Dongyoung, et al.
Pubblicazione: (2025)
di: Park, Dongyoung, et al.
Pubblicazione: (2025)
On the Variational Interpretation of Mirror Play in Monotone Games
di: Pan, Yunian, et al.
Pubblicazione: (2024)
di: Pan, Yunian, et al.
Pubblicazione: (2024)
Robust Deep Monte Carlo Counterfactual Regret Minimization: Addressing Theoretical Risks in Neural Fictitious Self-Play
di: Jaafari, Zakaria El
Pubblicazione: (2025)
di: Jaafari, Zakaria El
Pubblicazione: (2025)
Jailbreaking Large Language Models in Infinitely Many Ways
di: Goldstein, Oliver, et al.
Pubblicazione: (2025)
di: Goldstein, Oliver, et al.
Pubblicazione: (2025)
Mirror Mode in Fire Emblem: Beating Players at their own Game with Imitation and Reinforcement Learning
di: Smid, Yanna Elizabeth, et al.
Pubblicazione: (2025)
di: Smid, Yanna Elizabeth, et al.
Pubblicazione: (2025)
Analytical Stackelberg Resource Allocation in Sequential Attacker--Defender Games
di: Iqbal, Azhar, et al.
Pubblicazione: (2025)
di: Iqbal, Azhar, et al.
Pubblicazione: (2025)
Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games
di: Lanier, JB, et al.
Pubblicazione: (2026)
di: Lanier, JB, et al.
Pubblicazione: (2026)
Self-Play Q-learners Can Provably Collude in the Iterated Prisoner's Dilemma
di: Bertrand, Quentin, et al.
Pubblicazione: (2023)
di: Bertrand, Quentin, et al.
Pubblicazione: (2023)
A Scalable Communication Protocol for Networks of Large Language Models
di: Marro, Samuele, et al.
Pubblicazione: (2024)
di: Marro, Samuele, et al.
Pubblicazione: (2024)
Resolving social dilemmas with minimal reward transfer
di: Willis, Richard, et al.
Pubblicazione: (2023)
di: Willis, Richard, et al.
Pubblicazione: (2023)
Playing Large Games with Oracles and AI Debate
di: Chen, Xinyi, et al.
Pubblicazione: (2023)
di: Chen, Xinyi, et al.
Pubblicazione: (2023)
Playing games with Large language models: Randomness and strategy
di: Vidler, Alicia, et al.
Pubblicazione: (2025)
di: Vidler, Alicia, et al.
Pubblicazione: (2025)
Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play
di: Cipolina-Kun, Lucia, et al.
Pubblicazione: (2025)
di: Cipolina-Kun, Lucia, et al.
Pubblicazione: (2025)
Integrated Resource Allocation and Strategy Synthesis in Safety Games on Graphs with Deception
di: Kulkarni, Abhishek N., et al.
Pubblicazione: (2024)
di: Kulkarni, Abhishek N., et al.
Pubblicazione: (2024)
Study and improvement of search algorithms in two-players perfect information games
di: Cohen-Solal, Quentin
Pubblicazione: (2025)
di: Cohen-Solal, Quentin
Pubblicazione: (2025)
Study and Improvement of Search Algorithms in Multi-Player Perfect-Information Games
di: Cohen-Solal, Quentin
Pubblicazione: (2026)
di: Cohen-Solal, Quentin
Pubblicazione: (2026)
Large Language Models Playing Mixed Strategy Nash Equilibrium Games
di: Silva, Alonso
Pubblicazione: (2024)
di: Silva, Alonso
Pubblicazione: (2024)
Documenti analoghi
-
Tacit Coordination of Large Language Models
di: Aharon, Ido, et al.
Pubblicazione: (2026) -
Large Language Models Miss the Multi-Agent Mark
di: La Malfa, Emanuele, et al.
Pubblicazione: (2025) -
Using Protected Attributes to Consider Fairness in Multi-Agent Systems
di: La Malfa, Gabriele, et al.
Pubblicazione: (2024) -
Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI
di: Huang, Xuanqiang Angelo, et al.
Pubblicazione: (2026) -
Deep Neural Networks via Complex Network Theory: a Perspective
di: La Malfa, Emanuele, et al.
Pubblicazione: (2024)