SafeRBench: Dissecting the Reasoning Safety of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Xin, Yu, Shaohan, Chen, Zerui, Lyu, Yueming, Yu, Weichen, Li, Guanghao, Liu, Jiyao, Gao, Jianxiong, Liang, Jian, Liu, Ziwei, Si, Chenyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911398694486016
author Gao, Xin
Yu, Shaohan
Chen, Zerui
Lyu, Yueming
Yu, Weichen
Li, Guanghao
Liu, Jiyao
Gao, Jianxiong
Liang, Jian
Liu, Ziwei
Si, Chenyang
author_facet Gao, Xin
Yu, Shaohan
Chen, Zerui
Lyu, Yueming
Yu, Weichen
Li, Guanghao
Liu, Jiyao
Gao, Jianxiong
Liang, Jian
Liu, Ziwei
Si, Chenyang
contents Large Reasoning Models (LRMs) have significantly improved problem-solving through explicit Chain-of-Thought (CoT) reasoning. However, this capability creates a Safety-Helpfulness Paradox: the reasoning process itself can be misused to justify harmful actions or conceal malicious intent behind lengthy intermediate steps. Most existing benchmarks only check the final output, missing how risks evolve, or ``drift'', during the model's internal reasoning. To address this, we propose SafeRBench, the first framework to evaluate LRM safety end-to-end, from the initial input to the reasoning trace and final answer. Our approach introduces: (i) a Risk Stratification Probing that uses specific risk levels to stress-test safety boundaries beyond simple topics; (ii) Micro-Thought Analysis, a new chunking method that segments traces to pinpoint exactly where safety alignment breaks down; and (iii) a comprehensive suite of 10 fine-grained metrics that, for the first time, jointly measure a model's Risk Exposure (e.g., risk level, execution feasibility) and Safety Awareness (e.g., intent awareness). Experiments on 19 LRMs reveal that while enabling Thinking modes improves safety in mid-sized models, it paradoxically increases actionable risks in larger models due to a strong always-help tendency.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15169
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeRBench: Dissecting the Reasoning Safety of Large Language Models
Gao, Xin
Yu, Shaohan
Chen, Zerui
Lyu, Yueming
Yu, Weichen
Li, Guanghao
Liu, Jiyao
Gao, Jianxiong
Liang, Jian
Liu, Ziwei
Si, Chenyang
Artificial Intelligence
Large Reasoning Models (LRMs) have significantly improved problem-solving through explicit Chain-of-Thought (CoT) reasoning. However, this capability creates a Safety-Helpfulness Paradox: the reasoning process itself can be misused to justify harmful actions or conceal malicious intent behind lengthy intermediate steps. Most existing benchmarks only check the final output, missing how risks evolve, or ``drift'', during the model's internal reasoning. To address this, we propose SafeRBench, the first framework to evaluate LRM safety end-to-end, from the initial input to the reasoning trace and final answer. Our approach introduces: (i) a Risk Stratification Probing that uses specific risk levels to stress-test safety boundaries beyond simple topics; (ii) Micro-Thought Analysis, a new chunking method that segments traces to pinpoint exactly where safety alignment breaks down; and (iii) a comprehensive suite of 10 fine-grained metrics that, for the first time, jointly measure a model's Risk Exposure (e.g., risk level, execution feasibility) and Safety Awareness (e.g., intent awareness). Experiments on 19 LRMs reveal that while enabling Thinking modes improves safety in mid-sized models, it paradoxically increases actionable risks in larger models due to a strong always-help tendency.
title SafeRBench: Dissecting the Reasoning Safety of Large Language Models
topic Artificial Intelligence
url https://arxiv.org/abs/2511.15169