MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Weiyang, Li, Jing, Wang, Wenya, LI, YU, He, Daojing, Yu, Jun, Zhang, Min
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916754307940352
author Guo, Weiyang
Li, Jing
Wang, Wenya
LI, YU
He, Daojing
Yu, Jun
Zhang, Min
author_facet Guo, Weiyang
Li, Jing
Wang, Wenya
LI, YU
He, Daojing
Yu, Jun
Zhang, Min
contents The proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures. However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more prone to produce harmful responses. In this paper, we propose the \textbf{M}ulti-\textbf{T}urn \textbf{S}afety \textbf{A}lignment (\ourapproach) framework, to address the challenge of securing LLMs in multi-round interactions. It consists of two stages: In the thought-guided attack learning stage, the red-team model learns about thought-guided multi-round jailbreak attacks to generate adversarial prompts. In the adversarial iterative optimization stage, the red-team model and the target model continuously improve their respective capabilities in interaction. Furthermore, we introduce a multi-turn reinforcement learning algorithm based on future rewards to enhance the robustness of safety alignment. Experimental results show that the red-team model exhibits state-of-the-art attack capabilities, while the target model significantly improves its performance on safety benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17147
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
Guo, Weiyang
Li, Jing
Wang, Wenya
LI, YU
He, Daojing
Yu, Jun
Zhang, Min
Cryptography and Security
Artificial Intelligence
The proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures. However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more prone to produce harmful responses. In this paper, we propose the \textbf{M}ulti-\textbf{T}urn \textbf{S}afety \textbf{A}lignment (\ourapproach) framework, to address the challenge of securing LLMs in multi-round interactions. It consists of two stages: In the thought-guided attack learning stage, the red-team model learns about thought-guided multi-round jailbreak attacks to generate adversarial prompts. In the adversarial iterative optimization stage, the red-team model and the target model continuously improve their respective capabilities in interaction. Furthermore, we introduce a multi-turn reinforcement learning algorithm based on future rewards to enhance the robustness of safety alignment. Experimental results show that the red-team model exhibits state-of-the-art attack capabilities, while the target model significantly improves its performance on safety benchmarks.
title MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2505.17147