R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: In, Yeonjun, Kim, Wonjoong, Park, Sangwu, Park, Chanyoung
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912514953969664
author In, Yeonjun
Kim, Wonjoong
Park, Sangwu
Park, Chanyoung
author_facet In, Yeonjun
Kim, Wonjoong
Park, Sangwu
Park, Chanyoung
contents Although large reasoning models (LRMs) have demonstrated impressive capabilities on complex tasks, recent studies reveal that these models frequently fulfill harmful user instructions, raising significant safety concerns. In this paper, we investigate the underlying cause of LRM safety risks and find that models already possess sufficient safety knowledge but fail to activate it during reasoning. Based on this insight, we propose R1-Act, a simple and efficient post-training method that explicitly triggers safety knowledge through a structured reasoning process. R1-Act achieves strong safety improvements while preserving reasoning performance, outperforming prior alignment methods. Notably, it requires only 1,000 training examples and 90 minutes of training on a single RTX A6000 GPU. Extensive experiments across multiple LRM backbones and sizes demonstrate the robustness, scalability, and practical efficiency of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00324
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
In, Yeonjun
Kim, Wonjoong
Park, Sangwu
Park, Chanyoung
Artificial Intelligence
Computation and Language
Although large reasoning models (LRMs) have demonstrated impressive capabilities on complex tasks, recent studies reveal that these models frequently fulfill harmful user instructions, raising significant safety concerns. In this paper, we investigate the underlying cause of LRM safety risks and find that models already possess sufficient safety knowledge but fail to activate it during reasoning. Based on this insight, we propose R1-Act, a simple and efficient post-training method that explicitly triggers safety knowledge through a structured reasoning process. R1-Act achieves strong safety improvements while preserving reasoning performance, outperforming prior alignment methods. Notably, it requires only 1,000 training examples and 90 minutes of training on a single RTX A6000 GPU. Extensive experiments across multiple LRM backbones and sizes demonstrate the robustness, scalability, and practical efficiency of our approach.
title R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.00324