STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wu, Di, Zhao, Yanyan, Lu, Xin, Li, Mingzhe, Qin, Bing
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917187795550208
author Wu, Di
Zhao, Yanyan
Lu, Xin
Li, Mingzhe
Qin, Bing
author_facet Wu, Di
Zhao, Yanyan
Lu, Xin
Li, Mingzhe
Qin, Bing
contents Defending against jailbreak attacks is crucial for the safe deployment of Large Language Models (LLMs). Recent research has attempted to improve safety by training models to reason over safety rules before responding. However, a key issue lies in determining what form of safety reasoning effectively defends against jailbreak attacks, which is difficult to explicitly design or directly obtain. To address this, we propose \textbf{STAR-S} (\textbf{S}elf-\textbf{TA}ught \textbf{R}easoning based on \textbf{S}afety rules), a framework that integrates the learning of safety rule reasoning into a self-taught loop. The core of STAR-S involves eliciting reasoning and reflection guided by safety rules, then leveraging fine-tuning to enhance safety reasoning. Repeating this process creates a synergistic cycle. Improvements in the model's reasoning and interpretation of safety rules allow it to produce better reasoning data under safety rule prompts, which is then utilized for further training. Experiments show that STAR-S effectively defends against jailbreak attacks, outperforming baselines. Code is available at: https://github.com/pikepokenew/STAR_S.git.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03537
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
Wu, Di
Zhao, Yanyan
Lu, Xin
Li, Mingzhe
Qin, Bing
Artificial Intelligence
Computation and Language
Defending against jailbreak attacks is crucial for the safe deployment of Large Language Models (LLMs). Recent research has attempted to improve safety by training models to reason over safety rules before responding. However, a key issue lies in determining what form of safety reasoning effectively defends against jailbreak attacks, which is difficult to explicitly design or directly obtain. To address this, we propose \textbf{STAR-S} (\textbf{S}elf-\textbf{TA}ught \textbf{R}easoning based on \textbf{S}afety rules), a framework that integrates the learning of safety rule reasoning into a self-taught loop. The core of STAR-S involves eliciting reasoning and reflection guided by safety rules, then leveraging fine-tuning to enhance safety reasoning. Repeating this process creates a synergistic cycle. Improvements in the model's reasoning and interpretation of safety rules allow it to produce better reasoning data under safety rule prompts, which is then utilized for further training. Experiments show that STAR-S effectively defends against jailbreak attacks, outperforming baselines. Code is available at: https://github.com/pikepokenew/STAR_S.git.
title STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.03537