RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yichi, Zeng, Zihao, Li, Dongbai, Huang, Yao, Deng, Zhijie, Dong, Yinpeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913792429916160
author Zhang, Yichi
Zeng, Zihao
Li, Dongbai
Huang, Yao
Deng, Zhijie
Dong, Yinpeng
author_facet Zhang, Yichi
Zeng, Zihao
Li, Dongbai
Huang, Yao
Deng, Zhijie
Dong, Yinpeng
contents Large Reasoning Models (LRMs), such as OpenAI o1 and DeepSeek-R1, have been rapidly progressing and achieving breakthrough performance on complex reasoning tasks such as mathematics and coding. However, the open-source R1 models have raised safety concerns in wide applications, such as the tendency to comply with malicious queries, which greatly impacts the utility of these powerful models in their applications. In this paper, we introduce RealSafe-R1 as safety-aligned versions of DeepSeek-R1 distilled models. To train these models, we construct a dataset of 15k safety-aware reasoning trajectories generated by DeepSeek-R1, under explicit instructions for expected refusal behavior. Both quantitative experiments and qualitative case studies demonstrate the models' improvements, which are shown in their safety guardrails against both harmful queries and jailbreak attacks. Importantly, unlike prior safety alignment efforts that often compromise reasoning performance, our method preserves the models' reasoning capabilities by maintaining the training data within the original distribution of generation. Model weights of RealSafe-R1 are open-source at https://huggingface.co/RealSafe.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10081
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
Zhang, Yichi
Zeng, Zihao
Li, Dongbai
Huang, Yao
Deng, Zhijie
Dong, Yinpeng
Artificial Intelligence
Computation and Language
Large Reasoning Models (LRMs), such as OpenAI o1 and DeepSeek-R1, have been rapidly progressing and achieving breakthrough performance on complex reasoning tasks such as mathematics and coding. However, the open-source R1 models have raised safety concerns in wide applications, such as the tendency to comply with malicious queries, which greatly impacts the utility of these powerful models in their applications. In this paper, we introduce RealSafe-R1 as safety-aligned versions of DeepSeek-R1 distilled models. To train these models, we construct a dataset of 15k safety-aware reasoning trajectories generated by DeepSeek-R1, under explicit instructions for expected refusal behavior. Both quantitative experiments and qualitative case studies demonstrate the models' improvements, which are shown in their safety guardrails against both harmful queries and jailbreak attacks. Importantly, unlike prior safety alignment efforts that often compromise reasoning performance, our method preserves the models' reasoning capabilities by maintaining the training data within the original distribution of generation. Model weights of RealSafe-R1 are open-source at https://huggingface.co/RealSafe.
title RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.10081