SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Xianya, Luo, Xianying, Wang, Yadong, Chen, Xiang, Tian, Yu, Sun, Zequn, Liu, Rui, Fang, Jun, Tan, Naiqiang, Cui, Yuanning, Huang, Sheng-Jun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning
by: Luo, Qin-Wen, et al.
Published: (2026)
by: Luo, Qin-Wen, et al.
Published: (2026)
Breaking the Reasoning Horizon in Entity Alignment Foundation Models
by: Cui, Yuanning, et al.
Published: (2026)
by: Cui, Yuanning, et al.
Published: (2026)
A Prompt-Based Knowledge Graph Foundation Model for Universal In-Context Reasoning
by: Cui, Yuanning, et al.
Published: (2024)
by: Cui, Yuanning, et al.
Published: (2024)
Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs
by: Fang, Xianya, et al.
Published: (2026)
by: Fang, Xianya, et al.
Published: (2026)
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
by: Ning, Yansong, et al.
Published: (2025)
by: Ning, Yansong, et al.
Published: (2025)
Agent-Omit: Adaptive Context Omission for Efficient LLM Agents
by: Ning, Yansong, et al.
Published: (2026)
by: Ning, Yansong, et al.
Published: (2026)
KGFR: A Foundation Retriever for Generalized Knowledge Graph Question Answering
by: Cui, Yuanning, et al.
Published: (2025)
by: Cui, Yuanning, et al.
Published: (2025)
DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents
by: Ning, Yansong, et al.
Published: (2025)
by: Ning, Yansong, et al.
Published: (2025)
STAIR: Improving Safety Alignment with Introspective Reasoning
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
by: Xie, Yingsha, et al.
Published: (2026)
by: Xie, Yingsha, et al.
Published: (2026)
VeriThinker: Learning to Verify Makes Reasoning Model Efficient
by: Chen, Zigeng, et al.
Published: (2025)
by: Chen, Zigeng, et al.
Published: (2025)
TypedThinker: Diversify Large Language Model Reasoning with Typed Thinking
by: Wang, Danqing, et al.
Published: (2024)
by: Wang, Danqing, et al.
Published: (2024)
DiMA: An LLM-Powered Ride-Hailing Assistant at DiDi
by: Ning, Yansong, et al.
Published: (2025)
by: Ning, Yansong, et al.
Published: (2025)
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
by: Fang, Junfeng, et al.
Published: (2025)
by: Fang, Junfeng, et al.
Published: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025)
by: Fang, I-Sheng, et al.
Published: (2025)
Bag of Tricks for Inference-time Computation of LLM Reasoning
by: Liu, Fan, et al.
Published: (2025)
by: Liu, Fan, et al.
Published: (2025)
Expanding the Scope: Inductive Knowledge Graph Reasoning with Multi-Starting Progressive Propagation
by: Shao, Zhoutian, et al.
Published: (2024)
by: Shao, Zhoutian, et al.
Published: (2024)
Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
by: Wang, Zhaoxin, et al.
Published: (2026)
by: Wang, Zhaoxin, et al.
Published: (2026)
Does Postgraduate Education Deepen Temporomandibular Disorders Insights for Dental Professionals?
by: Zejin Liu, et al.
Published: (2024)
by: Zejin Liu, et al.
Published: (2024)
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
by: Sun, Chenkai, et al.
Published: (2025)
by: Sun, Chenkai, et al.
Published: (2025)
Beyond Dense States: Elevating Sparse Transcoders to Active Operators for Latent Reasoning
by: Wang, Yadong, et al.
Published: (2026)
by: Wang, Yadong, et al.
Published: (2026)
Generating Explanations to Understand and Repair Embedding-based Entity Alignment
by: Tian, Xiaobin, et al.
Published: (2023)
by: Tian, Xiaobin, et al.
Published: (2023)
ReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence Control
by: Tang, Zhentao, et al.
Published: (2026)
by: Tang, Zhentao, et al.
Published: (2026)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
by: Zhao, Yuze, et al.
Published: (2026)
by: Zhao, Yuze, et al.
Published: (2026)
Unified Thinker: A General Reasoning Modular Core for Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
Deepening Divides
Published: (2020)
Published: (2020)
Molecular mechanism analyses of post‐traumatic epilepsy and hereditary epilepsy based on 10× single‐cell transcriptome sequencing technology
by: Fang Wen, et al.
Published: (2024)
by: Fang Wen, et al.
Published: (2024)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing
by: Yang, Hanqing, et al.
Published: (2026)
by: Yang, Hanqing, et al.
Published: (2026)
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
by: Zheng, Baihui, et al.
Published: (2025)
by: Zheng, Baihui, et al.
Published: (2025)
The Missing Half: Unveiling Training-time Implicit Safety Risks Beyond Deployment
by: Zhang, Zhexin, et al.
Published: (2026)
by: Zhang, Zhexin, et al.
Published: (2026)
KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
by: Zhang, Dalong, et al.
Published: (2025)
by: Zhang, Dalong, et al.
Published: (2025)
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents
by: Wang, Shengqin, et al.
Published: (2026)
by: Wang, Shengqin, et al.
Published: (2026)
SafeWorld: Geo-Diverse Safety Alignment
by: Yin, Da, et al.
Published: (2024)
by: Yin, Da, et al.
Published: (2024)
SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving
by: Tian, Kefei, et al.
Published: (2026)
by: Tian, Kefei, et al.
Published: (2026)
The Association Between Socioeconomic Status and Cardiovascular Disease Risk in American Adults: Construction and Validation of a Nomogram Prediction Model Based on LASSO Feature Selection
by: Jun Li, et al.
Published: (2026)
by: Jun Li, et al.
Published: (2026)
SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse World
by: Kim, Jungho, et al.
Published: (2026)
by: Kim, Jungho, et al.
Published: (2026)
The Thinker
Published: (2021)
Published: (2021)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
by: Ding, Shengyuan, et al.
Published: (2025)
by: Ding, Shengyuan, et al.
Published: (2025)
Similar Items
-
Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning
by: Luo, Qin-Wen, et al.
Published: (2026) -
Breaking the Reasoning Horizon in Entity Alignment Foundation Models
by: Cui, Yuanning, et al.
Published: (2026) -
A Prompt-Based Knowledge Graph Foundation Model for Universal In-Context Reasoning
by: Cui, Yuanning, et al.
Published: (2024) -
Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs
by: Fang, Xianya, et al.
Published: (2026) -
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
by: Ning, Yansong, et al.
Published: (2025)