R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912514953969664 |
|---|---|
| author | In, Yeonjun Kim, Wonjoong Park, Sangwu Park, Chanyoung |
| author_facet | In, Yeonjun Kim, Wonjoong Park, Sangwu Park, Chanyoung |
| contents | Although large reasoning models (LRMs) have demonstrated impressive capabilities on complex tasks, recent studies reveal that these models frequently fulfill harmful user instructions, raising significant safety concerns. In this paper, we investigate the underlying cause of LRM safety risks and find that models already possess sufficient safety knowledge but fail to activate it during reasoning. Based on this insight, we propose R1-Act, a simple and efficient post-training method that explicitly triggers safety knowledge through a structured reasoning process. R1-Act achieves strong safety improvements while preserving reasoning performance, outperforming prior alignment methods. Notably, it requires only 1,000 training examples and 90 minutes of training on a single RTX A6000 GPU. Extensive experiments across multiple LRM backbones and sizes demonstrate the robustness, scalability, and practical efficiency of our approach. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_00324 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge In, Yeonjun Kim, Wonjoong Park, Sangwu Park, Chanyoung Artificial Intelligence Computation and Language Although large reasoning models (LRMs) have demonstrated impressive capabilities on complex tasks, recent studies reveal that these models frequently fulfill harmful user instructions, raising significant safety concerns. In this paper, we investigate the underlying cause of LRM safety risks and find that models already possess sufficient safety knowledge but fail to activate it during reasoning. Based on this insight, we propose R1-Act, a simple and efficient post-training method that explicitly triggers safety knowledge through a structured reasoning process. R1-Act achieves strong safety improvements while preserving reasoning performance, outperforming prior alignment methods. Notably, it requires only 1,000 training examples and 90 minutes of training on a single RTX A6000 GPU. Extensive experiments across multiple LRM backbones and sizes demonstrate the robustness, scalability, and practical efficiency of our approach. |
| title | R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2508.00324 |