Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xie, Yingsha, Huang, Tiansheng, Yang, Enneng, Min, Rui, Lu, Wenjie, Cao, Xiaochun, Tan, Naiqiang, Shen, Li
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912868477173760
author Xie, Yingsha
Huang, Tiansheng
Yang, Enneng
Min, Rui
Lu, Wenjie
Cao, Xiaochun
Tan, Naiqiang
Shen, Li
author_facet Xie, Yingsha
Huang, Tiansheng
Yang, Enneng
Min, Rui
Lu, Wenjie
Cao, Xiaochun
Tan, Naiqiang
Shen, Li
contents Safety alignment incurs safety tax that perturbs a large reasoning model's (LRM) general reasoning ability. Existing datasets used for safety alignment for an LRM are usually constructed by distilling safety reasoning traces and answers from an external LRM or human labeler. However, such reasoning traces and answers exhibit a distributional gap with the target LRM that needs alignment, and we conjecture such distributional gap is the culprit leading to significant degradation of reasoning ability of the target LRM. Driven by this hypothesis, we propose a safety alignment dataset construction method, dubbed DGR. DGR transforms and refines an existing out-of-distributional safety reasoning dataset to be aligned with the target's LLM inner distribution. Experimental results demonstrate that i) DGR effectively mitigates the safety tax while maintaining safety performance across all baselines, i.e., achieving \textbf{+30.2\%} on DirectRefusal and \textbf{+21.2\%} on R1-ACT improvement in average reasoning accuracy compared to Vanilla SFT; ii) the degree of reasoning degradation correlates with the extent of distribution shift, suggesting that bridging this gap is central to preserving capabilities. Furthermore, we find that safety alignment in LRMs may primarily function as a mechanism to activate latent knowledge, as a mere \textbf{10} samples are sufficient for activating effective refusal behaviors. These findings not only emphasize the importance of distributional consistency but also provide insights into the activation mechanism of safety in reasoning models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02136
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
Xie, Yingsha
Huang, Tiansheng
Yang, Enneng
Min, Rui
Lu, Wenjie
Cao, Xiaochun
Tan, Naiqiang
Shen, Li
Artificial Intelligence
Safety alignment incurs safety tax that perturbs a large reasoning model's (LRM) general reasoning ability. Existing datasets used for safety alignment for an LRM are usually constructed by distilling safety reasoning traces and answers from an external LRM or human labeler. However, such reasoning traces and answers exhibit a distributional gap with the target LRM that needs alignment, and we conjecture such distributional gap is the culprit leading to significant degradation of reasoning ability of the target LRM. Driven by this hypothesis, we propose a safety alignment dataset construction method, dubbed DGR. DGR transforms and refines an existing out-of-distributional safety reasoning dataset to be aligned with the target's LLM inner distribution. Experimental results demonstrate that i) DGR effectively mitigates the safety tax while maintaining safety performance across all baselines, i.e., achieving \textbf{+30.2\%} on DirectRefusal and \textbf{+21.2\%} on R1-ACT improvement in average reasoning accuracy compared to Vanilla SFT; ii) the degree of reasoning degradation correlates with the extent of distribution shift, suggesting that bridging this gap is central to preserving capabilities. Furthermore, we find that safety alignment in LRMs may primarily function as a mechanism to activate latent knowledge, as a mere \textbf{10} samples are sufficient for activating effective refusal behaviors. These findings not only emphasize the importance of distributional consistency but also provide insights into the activation mechanism of safety in reasoning models.
title Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
topic Artificial Intelligence
url https://arxiv.org/abs/2602.02136