STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Di, Zhao, Yanyan, Lu, Xin, Li, Mingzhe, Qin, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
Self-Taught Evaluators
by: Wang, Tianlu, et al.
Published: (2024)
by: Wang, Tianlu, et al.
Published: (2024)
Safety Compliance: Rethinking LLM Safety Reasoning through the Lens of Compliance
by: Hu, Wenbin, et al.
Published: (2025)
by: Hu, Wenbin, et al.
Published: (2025)
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
by: In, Yeonjun, et al.
Published: (2025)
by: In, Yeonjun, et al.
Published: (2025)
Self-Taught Agentic Long Context Understanding
by: Zhuang, Yufan, et al.
Published: (2025)
by: Zhuang, Yufan, et al.
Published: (2025)
Safety Reasoning with Guidelines
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Cat-DPO: Category-Adaptive Safety Alignment
by: Yang, Tiankai, et al.
Published: (2026)
by: Yang, Tiankai, et al.
Published: (2026)
Safety Is Not Universal: The Selective Safety Trap in LLM Alignment
by: Brito, Iago Alves, et al.
Published: (2026)
by: Brito, Iago Alves, et al.
Published: (2026)
Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
by: Zelikman, Eric, et al.
Published: (2023)
by: Zelikman, Eric, et al.
Published: (2023)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
by: Zhu, Zihao, et al.
Published: (2025)
by: Zhu, Zihao, et al.
Published: (2025)
V-STaR: Training Verifiers for Self-Taught Reasoners
by: Hosseini, Arian, et al.
Published: (2024)
by: Hosseini, Arian, et al.
Published: (2024)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
by: Wang, Zijun, et al.
Published: (2025)
by: Wang, Zijun, et al.
Published: (2025)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
by: Sinha, Yash, et al.
Published: (2024)
by: Sinha, Yash, et al.
Published: (2024)
Safety Alignment via Constrained Knowledge Unlearning
by: Shi, Zesheng, et al.
Published: (2025)
by: Shi, Zesheng, et al.
Published: (2025)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
Lifelong Safety Alignment for Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
by: Mao, Yingzhi, et al.
Published: (2025)
by: Mao, Yingzhi, et al.
Published: (2025)
Towards Context-Invariant Safety Alignment for Large Language Models
by: Wang, Yixu, et al.
Published: (2026)
by: Wang, Yixu, et al.
Published: (2026)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
by: Alssum, Lama, et al.
Published: (2025)
by: Alssum, Lama, et al.
Published: (2025)
Improve Rule Retrieval and Reasoning with Self-Induction and Relevance ReEstimate
by: Huang, Ziyang, et al.
Published: (2025)
by: Huang, Ziyang, et al.
Published: (2025)
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
by: Koh, Woosung, et al.
Published: (2025)
by: Koh, Woosung, et al.
Published: (2025)
MESA: Improving MoE Safety Alignment via Decentralized Expertise
by: Sun, Yitong, et al.
Published: (2026)
by: Sun, Yitong, et al.
Published: (2026)
What Matters For Safety Alignment?
by: Li, Xing, et al.
Published: (2026)
by: Li, Xing, et al.
Published: (2026)
Advancing Large Language Model Attribution through Self-Improving
by: Huang, Lei, et al.
Published: (2024)
by: Huang, Lei, et al.
Published: (2024)
Large Language Models Are Self-Taught Reasoners: Enhancing LLM Applications via Tailored Problem-Solving Demonstrations
by: Ong, Kai Tzu-iunn, et al.
Published: (2024)
by: Ong, Kai Tzu-iunn, et al.
Published: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
SafeWorld: Geo-Diverse Safety Alignment
by: Yin, Da, et al.
Published: (2024)
by: Yin, Da, et al.
Published: (2024)
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Multimodal Situational Safety
by: Zhou, Kaiwen, et al.
Published: (2024)
by: Zhou, Kaiwen, et al.
Published: (2024)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
by: Li, Jianwei, et al.
Published: (2025)
by: Li, Jianwei, et al.
Published: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
by: Zhang, Zhehao, et al.
Published: (2025)
by: Zhang, Zhehao, et al.
Published: (2025)
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions
by: Xu, Jingxin, et al.
Published: (2025)
by: Xu, Jingxin, et al.
Published: (2025)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
by: Chen, Jianhui, et al.
Published: (2024)
by: Chen, Jianhui, et al.
Published: (2024)
ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models
by: Li, Changyi, et al.
Published: (2025)
by: Li, Changyi, et al.
Published: (2025)
Test-Time Safety Alignment
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
Similar Items
-
Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models
by: Wu, Di, et al.
Published: (2024) -
Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
by: Wang, Yuhang, et al.
Published: (2025) -
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026) -
Self-Taught Evaluators
by: Wang, Tianlu, et al.
Published: (2024) -
Safety Compliance: Rethinking LLM Safety Reasoning through the Lens of Compliance
by: Hu, Wenbin, et al.
Published: (2025)