When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Yingzhi, Zhang, Chunkang, Wang, Junxiang, Guan, Xinyan, Cao, Boxi, Lu, Yaojie, Lin, Hongyu, Han, Xianpei, Sun, Le
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913058654257152
author Mao, Yingzhi
Zhang, Chunkang
Wang, Junxiang
Guan, Xinyan
Cao, Boxi
Lu, Yaojie
Lin, Hongyu
Han, Xianpei
Sun, Le
author_facet Mao, Yingzhi
Zhang, Chunkang
Wang, Junxiang
Guan, Xinyan
Cao, Boxi
Lu, Yaojie
Lin, Hongyu
Han, Xianpei
Sun, Le
contents Large Reasoning Models (LRMs) achieve strong performance on complex multi-step reasoning, yet they still exhibit severe safety failures such as harmful content generation. Existing methods often apply coarse-grained constraints over the entire reasoning trajectories, which can undermine reasoning capability while failing to address the root causes of unsafe behavior. In this work, we uncover a previously underexplored failure mode in LRMs, termed Self-Jailbreak, where models initially recognize the harmful intent of a query, but override this judgment during subsequent reasoning steps, ultimately generating unsafe outputs. Such a phenomenon reveals that LRMs are capable of recognizing harm, while safety failures primarily arise from reasoning steps. Motivated by this finding, we propose Chain-of-Guardrail(CoG), a trajectory-level training framework that mitigates Self-Jailbreak via targeted, step-level interventions while maintaining reasoning ability. Experiments across multiple safety and reasoning benchmarks indicate that CoG achieves a favorable balance between safety and reasoning performance compared with existing approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21285
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
Mao, Yingzhi
Zhang, Chunkang
Wang, Junxiang
Guan, Xinyan
Cao, Boxi
Lu, Yaojie
Lin, Hongyu
Han, Xianpei
Sun, Le
Artificial Intelligence
Computation and Language
Large Reasoning Models (LRMs) achieve strong performance on complex multi-step reasoning, yet they still exhibit severe safety failures such as harmful content generation. Existing methods often apply coarse-grained constraints over the entire reasoning trajectories, which can undermine reasoning capability while failing to address the root causes of unsafe behavior. In this work, we uncover a previously underexplored failure mode in LRMs, termed Self-Jailbreak, where models initially recognize the harmful intent of a query, but override this judgment during subsequent reasoning steps, ultimately generating unsafe outputs. Such a phenomenon reveals that LRMs are capable of recognizing harm, while safety failures primarily arise from reasoning steps. Motivated by this finding, we propose Chain-of-Guardrail(CoG), a trajectory-level training framework that mitigates Self-Jailbreak via targeted, step-level interventions while maintaining reasoning ability. Experiments across multiple safety and reasoning benchmarks indicate that CoG achieves a favorable balance between safety and reasoning performance compared with existing approaches.
title When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.21285