TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiong, Xiqiao, Li, Ouxiang, Liu, Zhuo, Li, Moxin, Shi, Wentao, Zhu, Fengbin, Wang, Qifan, Feng, Fuli
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917424255729664
author Xiong, Xiqiao
Li, Ouxiang
Liu, Zhuo
Li, Moxin
Shi, Wentao
Zhu, Fengbin
Wang, Qifan
Feng, Fuli
author_facet Xiong, Xiqiao
Li, Ouxiang
Liu, Zhuo
Li, Moxin
Shi, Wentao
Zhu, Fengbin
Wang, Qifan
Feng, Fuli
contents Large language models have seen widespread adoption, yet they remain vulnerable to multi-turn jailbreak attacks, threatening their safe deployment. This has led to the task of training automated multi-turn attackers to probe model safety vulnerabilities. However, existing approaches typically rely on turn-level optimization, which is insufficient for learning long-term attack strategies. To bridge this gap, we formulate this task as a multi-turn reinforcement learning problem, directly optimizing the harmfulness of the final-turn response as the outcome reward. To address the sparse supervision of the outcome reward, we introduce TROJail, which employs two process rewards to evaluate the utility of intermediate prompts and integrate them into advantage estimation. These rewards (1) penalize overly harmful prompts that trigger the model's refusal mechanism, and (2) encourage steering the semantic relevance of responses toward the targeted harmful content. Experimental results show improved attack success rates across multiple models and benchmarks, highlighting the effectiveness of our approach. The code is available at https://github.com/xxiqiao/TROJail. Warning: This paper contains examples of harmful content.
format Preprint
id arxiv_https___arxiv_org_abs_2512_07761
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
Xiong, Xiqiao
Li, Ouxiang
Liu, Zhuo
Li, Moxin
Shi, Wentao
Zhu, Fengbin
Wang, Qifan
Feng, Fuli
Artificial Intelligence
Machine Learning
Large language models have seen widespread adoption, yet they remain vulnerable to multi-turn jailbreak attacks, threatening their safe deployment. This has led to the task of training automated multi-turn attackers to probe model safety vulnerabilities. However, existing approaches typically rely on turn-level optimization, which is insufficient for learning long-term attack strategies. To bridge this gap, we formulate this task as a multi-turn reinforcement learning problem, directly optimizing the harmfulness of the final-turn response as the outcome reward. To address the sparse supervision of the outcome reward, we introduce TROJail, which employs two process rewards to evaluate the utility of intermediate prompts and integrate them into advantage estimation. These rewards (1) penalize overly harmful prompts that trigger the model's refusal mechanism, and (2) encourage steering the semantic relevance of responses toward the targeted harmful content. Experimental results show improved attack success rates across multiple models and benchmarks, highlighting the effectiveness of our approach. The code is available at https://github.com/xxiqiao/TROJail. Warning: This paper contains examples of harmful content.
title TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.07761