TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Zhewen, Yu, Wenhan, Si, Jianfeng, Liu, Tongxin, Guan, Kaiqi, Jin, Huiyan, Tao, Jiawen, Yuan, Xiaokun, Ma, Duohe, Zhang, Xiangzheng, Yang, Tong, Sun, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
by: Si, Jianfeng, et al.
Published: (2025)
by: Si, Jianfeng, et al.
Published: (2025)
MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection
by: Tan, Zhewen, et al.
Published: (2026)
by: Tan, Zhewen, et al.
Published: (2026)
Beyond Static Alignment: Hierarchical Policy Control for LLM Safety via Risk-Aware Chain-of-Thought
by: Si, Jianfeng, et al.
Published: (2026)
by: Si, Jianfeng, et al.
Published: (2026)
Plug-and-Play Tri-Branch Invertible Block for Image Rescaling
by: Bao, Jingwei, et al.
Published: (2024)
by: Bao, Jingwei, et al.
Published: (2024)
FedSelfPlay-Defense: A Federated Adversarial Self-Play Framework for LLM Safety
by: Jayathilaka, Hasini
Published: (2025)
by: Jayathilaka, Hasini
Published: (2025)
Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment
by: Li, Jiajia, et al.
Published: (2026)
by: Li, Jiajia, et al.
Published: (2026)
SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
by: Fang, Wenkai, et al.
Published: (2025)
by: Fang, Wenkai, et al.
Published: (2025)
Try-On-Adapter: A Simple and Flexible Try-On Paradigm
by: Guo, Hanzhong, et al.
Published: (2024)
by: Guo, Hanzhong, et al.
Published: (2024)
Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment
by: Vega, Jason, et al.
Published: (2024)
by: Vega, Jason, et al.
Published: (2024)
Safe Exploitative Play with Untrusted Type Beliefs
by: Li, Tongxin, et al.
Published: (2024)
by: Li, Tongxin, et al.
Published: (2024)
PsyPlay: Personality-Infused Role-Playing Conversational Agents
by: Yang, Tao, et al.
Published: (2025)
by: Yang, Tao, et al.
Published: (2025)
TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment
by: Nguyen, Thi-Nhung, et al.
Published: (2026)
by: Nguyen, Thi-Nhung, et al.
Published: (2026)
SPARS: Self-Play Adversarial Reinforcement Learning for Segmentation of Liver Tumours
by: Tan, Catalina, et al.
Published: (2025)
by: Tan, Catalina, et al.
Published: (2025)
MOA: Multi-Objective Alignment for Role-Playing Agents
by: Liao, Chonghua, et al.
Published: (2025)
by: Liao, Chonghua, et al.
Published: (2025)
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play
by: Ye, Ziyu, et al.
Published: (2024)
by: Ye, Ziyu, et al.
Published: (2024)
CORAL: Correspondence Alignment for Improved Virtual Try-On
by: Kim, Jiyoung, et al.
Published: (2026)
by: Kim, Jiyoung, et al.
Published: (2026)
Self-Play Preference Optimization for Language Model Alignment
by: Wu, Yue, et al.
Published: (2024)
by: Wu, Yue, et al.
Published: (2024)
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
by: Xie, Yuanbo, et al.
Published: (2025)
by: Xie, Yuanbo, et al.
Published: (2025)
Self-Supervised Vision Transformer for Enhanced Virtual Clothes Try-On
by: Lu, Lingxiao, et al.
Published: (2024)
by: Lu, Lingxiao, et al.
Published: (2024)
Differentially Private Reinforcement Learning with Self-Play
by: Qiao, Dan, et al.
Published: (2024)
by: Qiao, Dan, et al.
Published: (2024)
Internalizing World Models via Self-Play Finetuning for Agentic RL
by: Chen, Shiqi, et al.
Published: (2025)
by: Chen, Shiqi, et al.
Published: (2025)
Towards Customized Multimodal Role-Play
by: Tang, Chao, et al.
Published: (2026)
by: Tang, Chao, et al.
Published: (2026)
Tri-Level Navigator: LLM-Empowered Tri-Level Learning for Time Series OOD Generalization
by: Jian, Chengtao, et al.
Published: (2024)
by: Jian, Chengtao, et al.
Published: (2024)
Play to Earn in the Metaverse with Mobile Edge Computing over Wireless Networks: A Deep Reinforcement Learning Approach
by: Chua, Terence Jie, et al.
Published: (2023)
by: Chua, Terence Jie, et al.
Published: (2023)
MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
RSPO: Regularized Self-Play Alignment of Large Language Models
by: Tang, Xiaohang, et al.
Published: (2025)
by: Tang, Xiaohang, et al.
Published: (2025)
From Role-Play to Drama-Interaction: An LLM Solution
by: Wu, Weiqi, et al.
Published: (2024)
by: Wu, Weiqi, et al.
Published: (2024)
RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents
by: Lai, Huayi, et al.
Published: (2026)
by: Lai, Huayi, et al.
Published: (2026)
Library Role-Playing.
by: Chapman, Liz
Published: (1979)
by: Chapman, Liz
Published: (1979)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Reproducing AlphaZero on Tablut: Self-Play RL for an Asymmetric Board Game
by: Lees, Tõnis, et al.
Published: (2026)
by: Lees, Tõnis, et al.
Published: (2026)
OmniTry: Virtual Try-On Anything without Masks
by: Feng, Yutong, et al.
Published: (2025)
by: Feng, Yutong, et al.
Published: (2025)
If at First You Don't Succeed, Try (and Try) Again
Published: (2025)
Published: (2025)
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
by: Li, Shangzhe, et al.
Published: (2026)
by: Li, Shangzhe, et al.
Published: (2026)
SPRec: Self-Play to Debias LLM-based Recommendation
by: Gao, Chongming, et al.
Published: (2024)
by: Gao, Chongming, et al.
Published: (2024)
Concept Incongruence: An Exploration of Time and Death in Role Playing
by: Bai, Xiaoyan, et al.
Published: (2025)
by: Bai, Xiaoyan, et al.
Published: (2025)
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing
by: Wang, Miao, et al.
Published: (2026)
by: Wang, Miao, et al.
Published: (2026)
Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
Stay in Character, Stay Safe: Dual-Cycle Adversarial Self-Evolution for Safety Role-Playing Agents
by: Liao, Mingyang, et al.
Published: (2026)
by: Liao, Mingyang, et al.
Published: (2026)
Similar Items
-
Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
by: Si, Jianfeng, et al.
Published: (2025) -
MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection
by: Tan, Zhewen, et al.
Published: (2026) -
Beyond Static Alignment: Hierarchical Policy Control for LLM Safety via Risk-Aware Chain-of-Thought
by: Si, Jianfeng, et al.
Published: (2026) -
Plug-and-Play Tri-Branch Invertible Block for Image Rescaling
by: Bao, Jingwei, et al.
Published: (2024) -
FedSelfPlay-Defense: A Federated Adversarial Self-Play Framework for LLM Safety
by: Jayathilaka, Hasini
Published: (2025)