TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Zhewen, Yu, Wenhan, Si, Jianfeng, Liu, Tongxin, Guan, Kaiqi, Jin, Huiyan, Tao, Jiawen, Yuan, Xiaokun, Ma, Duohe, Zhang, Xiangzheng, Yang, Tong, Sun, Lin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918315385946112
author Tan, Zhewen
Yu, Wenhan
Si, Jianfeng
Liu, Tongxin
Guan, Kaiqi
Jin, Huiyan
Tao, Jiawen
Yuan, Xiaokun
Ma, Duohe
Zhang, Xiangzheng
Yang, Tong
Sun, Lin
author_facet Tan, Zhewen
Yu, Wenhan
Si, Jianfeng
Liu, Tongxin
Guan, Kaiqi
Jin, Huiyan
Tao, Jiawen
Yuan, Xiaokun
Ma, Duohe
Zhang, Xiangzheng
Yang, Tong
Sun, Lin
contents In recent years, safety risks associated with large language models have become increasingly prominent, highlighting the urgent need to mitigate the generation of toxic and harmful content. The mainstream paradigm for LLM safety alignment typically adopts a collaborative framework involving three roles: an attacker for adversarial prompt generation, a defender for safety defense, and an evaluator for response assessment. In this paper, we propose a closed-loop reinforcement learning framework called TriPlay-RL that enables iterative and co-improving collaboration among three roles with near-zero manual annotation. Experimental results show that the attacker preserves high output diversity while achieving a 20%-50% improvement in adversarial effectiveness; the defender attains 10%-30% gains in safety performance without degrading general reasoning capability; and the evaluator continuously refines its fine-grained judgment ability through iterations, accurately distinguishing unsafe responses, simple refusals, and useful guidance. Overall, our framework establishes an efficient and scalable paradigm for LLM safety alignment, enabling continuous co-evolution within a unified learning loop.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18292
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
Tan, Zhewen
Yu, Wenhan
Si, Jianfeng
Liu, Tongxin
Guan, Kaiqi
Jin, Huiyan
Tao, Jiawen
Yuan, Xiaokun
Ma, Duohe
Zhang, Xiangzheng
Yang, Tong
Sun, Lin
Machine Learning
Artificial Intelligence
In recent years, safety risks associated with large language models have become increasingly prominent, highlighting the urgent need to mitigate the generation of toxic and harmful content. The mainstream paradigm for LLM safety alignment typically adopts a collaborative framework involving three roles: an attacker for adversarial prompt generation, a defender for safety defense, and an evaluator for response assessment. In this paper, we propose a closed-loop reinforcement learning framework called TriPlay-RL that enables iterative and co-improving collaboration among three roles with near-zero manual annotation. Experimental results show that the attacker preserves high output diversity while achieving a 20%-50% improvement in adversarial effectiveness; the defender attains 10%-30% gains in safety performance without degrading general reasoning capability; and the evaluator continuously refines its fine-grained judgment ability through iterations, accurately distinguishing unsafe responses, simple refusals, and useful guidance. Overall, our framework establishes an efficient and scalable paradigm for LLM safety alignment, enabling continuous co-evolution within a unified learning loop.
title TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.18292