SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Fengqing, Xu, Zhangchen, Li, Yuetai, Niu, Luyao, Xiang, Zhen, Li, Bo, Lin, Bill Yuchen, Poovendran, Radha
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912234068770816
author Jiang, Fengqing
Xu, Zhangchen
Li, Yuetai
Niu, Luyao
Xiang, Zhen
Li, Bo
Lin, Bill Yuchen
Poovendran, Radha
author_facet Jiang, Fengqing
Xu, Zhangchen
Li, Yuetai
Niu, Luyao
Xiang, Zhen
Li, Bo
Lin, Bill Yuchen
Poovendran, Radha
contents Emerging large reasoning models (LRMs), such as DeepSeek-R1 models, leverage long chain-of-thought (CoT) reasoning to generate structured intermediate steps, enhancing their reasoning capabilities. However, long CoT does not inherently guarantee safe outputs, potentially leading to harmful consequences such as the introduction of security vulnerabilities in code or the spread of misinformation. Current research on large language model (LLM) safety usually focuses on short-answer responses, overlooking the long CoT style outputs of LRMs. To bridge this gap, we conduct a systematic study of LRM safety. First, we investigate safety evaluators calibrated against human annotations. Using our newly developed metrics, we thoroughly assess the safety of 12 state-of-the-art LRMs on StrongReject and WildJailbreak datasets. Our results show that LRMs are not safe compared to their reasoning advance. Further, we perform a fine-grained analysis of the reasoning trace and final answer. We find that three decoding strategies-ZeroThink, LessThink, and MoreThink-can improve model safety without additional training. However, these strategies either use constrained reasoning traces or incur high inference costs. To better strengthen LRM safety, we introduce SafeChain, the first-of-its-kind safety training dataset in CoT style. We fine-tune two LRMs with SafeChain, showing that it not only enhances model safety but also preserves performance across 6 reasoning benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12025
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
Jiang, Fengqing
Xu, Zhangchen
Li, Yuetai
Niu, Luyao
Xiang, Zhen
Li, Bo
Lin, Bill Yuchen
Poovendran, Radha
Artificial Intelligence
Computation and Language
Emerging large reasoning models (LRMs), such as DeepSeek-R1 models, leverage long chain-of-thought (CoT) reasoning to generate structured intermediate steps, enhancing their reasoning capabilities. However, long CoT does not inherently guarantee safe outputs, potentially leading to harmful consequences such as the introduction of security vulnerabilities in code or the spread of misinformation. Current research on large language model (LLM) safety usually focuses on short-answer responses, overlooking the long CoT style outputs of LRMs. To bridge this gap, we conduct a systematic study of LRM safety. First, we investigate safety evaluators calibrated against human annotations. Using our newly developed metrics, we thoroughly assess the safety of 12 state-of-the-art LRMs on StrongReject and WildJailbreak datasets. Our results show that LRMs are not safe compared to their reasoning advance. Further, we perform a fine-grained analysis of the reasoning trace and final answer. We find that three decoding strategies-ZeroThink, LessThink, and MoreThink-can improve model safety without additional training. However, these strategies either use constrained reasoning traces or incur high inference costs. To better strengthen LRM safety, we introduce SafeChain, the first-of-its-kind safety training dataset in CoT style. We fine-tune two LRMs with SafeChain, showing that it not only enhances model safety but also preserves performance across 6 reasoning benchmarks.
title SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.12025