Safety Alignment of LMs via Non-cooperative Games

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paulus, Anselm, Kulikov, Ilia, Amos, Brandon, Munos, Rémi, Evtimov, Ivan, Chaudhuri, Kamalika, Zharmagambetov, Arman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911737512460288
author Paulus, Anselm
Kulikov, Ilia
Amos, Brandon
Munos, Rémi
Evtimov, Ivan
Chaudhuri, Kamalika
Zharmagambetov, Arman
author_facet Paulus, Anselm
Kulikov, Ilia
Amos, Brandon
Munos, Rémi
Evtimov, Ivan
Chaudhuri, Kamalika
Zharmagambetov, Arman
contents Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment. Current approaches rely on sequential adversarial training: generating adversarial prompts and fine-tuning LMs to defend against them. We introduce a different paradigm: framing safety alignment as a non-zero-sum game between an Attacker LM and a Defender LM trained jointly via online reinforcement learning. Each LM continuously adapts to the other's evolving strategies, driving iterative improvement. Our method uses a preference-based reward signal derived from pairwise comparisons instead of point-wise scores, providing more robust supervision and potentially reducing reward hacking. Our RL recipe, AdvGame, shifts the Pareto frontier of safety and utility, yielding a Defender LM that is simultaneously more helpful and more resilient to adversarial attacks. In addition, the resulting Attacker LM converges into a strong, general-purpose red-teaming agent that can be directly deployed to probe arbitrary target models. Code at github.com/facebookresearch/advgame.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Safety Alignment of LMs via Non-cooperative Games
Paulus, Anselm
Kulikov, Ilia
Amos, Brandon
Munos, Rémi
Evtimov, Ivan
Chaudhuri, Kamalika
Zharmagambetov, Arman
Artificial Intelligence
Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment. Current approaches rely on sequential adversarial training: generating adversarial prompts and fine-tuning LMs to defend against them. We introduce a different paradigm: framing safety alignment as a non-zero-sum game between an Attacker LM and a Defender LM trained jointly via online reinforcement learning. Each LM continuously adapts to the other's evolving strategies, driving iterative improvement. Our method uses a preference-based reward signal derived from pairwise comparisons instead of point-wise scores, providing more robust supervision and potentially reducing reward hacking. Our RL recipe, AdvGame, shifts the Pareto frontier of safety and utility, yielding a Defender LM that is simultaneously more helpful and more resilient to adversarial attacks. In addition, the resulting Attacker LM converges into a strong, general-purpose red-teaming agent that can be directly deployed to probe arbitrary target models. Code at github.com/facebookresearch/advgame.
title Safety Alignment of LMs via Non-cooperative Games
topic Artificial Intelligence
url https://arxiv.org/abs/2512.20806