Saved in:
Bibliographic Details
Main Authors: Lu, Chengda, Fan, Xiaoyu, Huang, Yu, Xu, Rongwu, Li, Jijie, Xu, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.17650
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Jailbreak attacks have been observed to largely fail against recent reasoning models enhanced by Chain-of-Thought (CoT) reasoning. However, the underlying mechanism remains underexplored, and relying solely on reasoning capacity may raise security concerns. In this paper, we try to answer the question: Does CoT reasoning really reduce harmfulness from jailbreaking? Through rigorous theoretical analysis, we demonstrate that CoT reasoning has dual effects on jailbreaking harmfulness. Based on the theoretical insights, we propose a novel jailbreak method, FicDetail, whose practical performance validates our theoretical findings.