Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Tiansheng, Hu, Sihao, Ilhan, Fatih, Tekin, Selim Furkan, Yahn, Zachary, Xu, Yichang, Liu, Ling
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913876230012928
author Huang, Tiansheng
Hu, Sihao
Ilhan, Fatih
Tekin, Selim Furkan
Yahn, Zachary
Xu, Yichang
Liu, Ling
author_facet Huang, Tiansheng
Hu, Sihao
Ilhan, Fatih
Tekin, Selim Furkan
Yahn, Zachary
Xu, Yichang
Liu, Ling
contents Safety alignment is an important procedure before the official deployment of a Large Language Model (LLM). While safety alignment has been extensively studied for LLM, there is still a large research gap for Large Reasoning Models (LRMs) that equip with improved reasoning capability. We in this paper systematically examine a simplified pipeline for producing safety aligned LRMs. With our evaluation of various LRMs, we deliver two main findings: i) Safety alignment can be done upon the LRM to restore its safety capability. ii) Safety alignment leads to a degradation of the reasoning capability of LRMs. The two findings show that there exists a trade-off between reasoning and safety capability with the sequential LRM production pipeline. The discovered trade-off, which we name Safety Tax, should shed light on future endeavors of safety research on LRMs. As a by-product, we curate a dataset called DirectRefusal, which might serve as an alternative dataset for safety alignment. Our source code is available at https://github.com/git-disl/Safety-Tax.
format Preprint
id arxiv_https___arxiv_org_abs_2503_00555
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
Huang, Tiansheng
Hu, Sihao
Ilhan, Fatih
Tekin, Selim Furkan
Yahn, Zachary
Xu, Yichang
Liu, Ling
Cryptography and Security
Artificial Intelligence
Machine Learning
Safety alignment is an important procedure before the official deployment of a Large Language Model (LLM). While safety alignment has been extensively studied for LLM, there is still a large research gap for Large Reasoning Models (LRMs) that equip with improved reasoning capability. We in this paper systematically examine a simplified pipeline for producing safety aligned LRMs. With our evaluation of various LRMs, we deliver two main findings: i) Safety alignment can be done upon the LRM to restore its safety capability. ii) Safety alignment leads to a degradation of the reasoning capability of LRMs. The two findings show that there exists a trade-off between reasoning and safety capability with the sequential LRM production pipeline. The discovered trade-off, which we name Safety Tax, should shed light on future endeavors of safety research on LRMs. As a by-product, we curate a dataset called DirectRefusal, which might serve as an alternative dataset for safety alignment. Our source code is available at https://github.com/git-disl/Safety-Tax.
title Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.00555