Internalizing Safety Understanding in Large Reasoning Models via Verification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yi, Chen, Yuxin, Sheng, Leheng, Zhang, Dongcheng, Lu, Chaochao, Wang, Xiang, Zhang, An |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
von: Zhang, Dongcheng, et al.
Veröffentlicht: (2026)
von: Zhang, Dongcheng, et al.
Veröffentlicht: (2026)
Reasoning Can Be Restored by Correcting a Few Decision Tokens
von: Shen, Changshuo, et al.
Veröffentlicht: (2026)
von: Shen, Changshuo, et al.
Veröffentlicht: (2026)
On Reasoning Strength Planning in Large Reasoning Models
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
von: Sheng, Leheng, et al.
Veröffentlicht: (2025)
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
Language Representations Can be What Recommenders Need: Findings and Potentials
von: Sheng, Leheng, et al.
Veröffentlicht: (2024)
von: Sheng, Leheng, et al.
Veröffentlicht: (2024)
On Generative Agents in Recommendation
von: Zhang, An, et al.
Veröffentlicht: (2023)
von: Zhang, An, et al.
Veröffentlicht: (2023)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
On Softmax Direct Preference Optimization for Recommendation
von: Chen, Yuxin, et al.
Veröffentlicht: (2024)
von: Chen, Yuxin, et al.
Veröffentlicht: (2024)
MiniOneRec: An Open-Source Framework for Scaling Generative Recommendation
von: Kong, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Kong, Xiaoyu, et al.
Veröffentlicht: (2025)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
von: Ding, Chenlu, et al.
Veröffentlicht: (2025)
von: Ding, Chenlu, et al.
Veröffentlicht: (2025)
SafeCoT: Improving VLM Safety with Minimal Reasoning
von: Ma, Jiachen, et al.
Veröffentlicht: (2025)
von: Ma, Jiachen, et al.
Veröffentlicht: (2025)
Risky-Bench: Probing Agentic Safety Risks under Real-World Deployment
von: Zheng, Jingnan, et al.
Veröffentlicht: (2026)
von: Zheng, Jingnan, et al.
Veröffentlicht: (2026)
Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
von: Yang, Junyao, et al.
Veröffentlicht: (2026)
von: Yang, Junyao, et al.
Veröffentlicht: (2026)
When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
Customizing Language Models with Instance-wise LoRA for Sequential Recommendation
von: Kong, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Kong, Xiaoyu, et al.
Veröffentlicht: (2024)
Owner-Harm: A Missing Threat Model for AI Agent Safety
von: Zhang, Dongcheng, et al.
Veröffentlicht: (2026)
von: Zhang, Dongcheng, et al.
Veröffentlicht: (2026)
KALE: Enhancing Knowledge Manipulation in Large Language Models via Knowledge-aware Learning
von: Lv, Qitan, et al.
Veröffentlicht: (2026)
von: Lv, Qitan, et al.
Veröffentlicht: (2026)
Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models
von: Wang, Wanying, et al.
Veröffentlicht: (2025)
von: Wang, Wanying, et al.
Veröffentlicht: (2025)
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak
von: Ma, Jiachen, et al.
Veröffentlicht: (2026)
von: Ma, Jiachen, et al.
Veröffentlicht: (2026)
Distribution-consistency Structural Causal Models
von: Gong, Heyang, et al.
Veröffentlicht: (2024)
von: Gong, Heyang, et al.
Veröffentlicht: (2024)
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
von: Chen, Benteng, et al.
Veröffentlicht: (2026)
von: Chen, Benteng, et al.
Veröffentlicht: (2026)
Understanding Chain-of-Thought in Large Language Models via Topological Data Analysis
von: Li, Chenghao, et al.
Veröffentlicht: (2025)
von: Li, Chenghao, et al.
Veröffentlicht: (2025)
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
von: Wang, Yuanfu, et al.
Veröffentlicht: (2026)
von: Wang, Yuanfu, et al.
Veröffentlicht: (2026)
Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning
von: Wang, Li, et al.
Veröffentlicht: (2025)
von: Wang, Li, et al.
Veröffentlicht: (2025)
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
von: Fang, Junfeng, et al.
Veröffentlicht: (2025)
von: Fang, Junfeng, et al.
Veröffentlicht: (2025)
Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training
von: Wang, Chen, et al.
Veröffentlicht: (2026)
von: Wang, Chen, et al.
Veröffentlicht: (2026)
Circular Reasoning: Understanding Self-Reinforcing Loops in Large Reasoning Models
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering
von: Li, Xiaomin, et al.
Veröffentlicht: (2026)
von: Li, Xiaomin, et al.
Veröffentlicht: (2026)
Internal Safety Collapse in Frontier Large Language Models
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
Benchmarking MLLM-based Web Understanding: Reasoning, Robustness and Safety
von: Liu, Junliang, et al.
Veröffentlicht: (2025)
von: Liu, Junliang, et al.
Veröffentlicht: (2025)
Towards Safer Large Reasoning Models by Promoting Safety Decision-Making before Chain-of-Thought Generation
von: Chen, Jianan, et al.
Veröffentlicht: (2026)
von: Chen, Jianan, et al.
Veröffentlicht: (2026)
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
von: Han, Bing, et al.
Veröffentlicht: (2025)
von: Han, Bing, et al.
Veröffentlicht: (2025)
A Closer Look at the Self-Verification Abilities of Large Language Models in Logical Reasoning
von: Hong, Ruixin, et al.
Veröffentlicht: (2023)
von: Hong, Ruixin, et al.
Veröffentlicht: (2023)
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
von: Chen, Sirui, et al.
Veröffentlicht: (2026)
von: Chen, Sirui, et al.
Veröffentlicht: (2026)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models
von: Zhang, Nan, et al.
Veröffentlicht: (2025)
von: Zhang, Nan, et al.
Veröffentlicht: (2025)
ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search
von: Zhang, Yize, et al.
Veröffentlicht: (2025)
von: Zhang, Yize, et al.
Veröffentlicht: (2025)
Dynamic Adversarial Reinforcement Learning for Robust Multimodal Large Language Models
von: Bao, Yicheng, et al.
Veröffentlicht: (2026)
von: Bao, Yicheng, et al.
Veröffentlicht: (2026)
Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
von: Xie, Yingsha, et al.
Veröffentlicht: (2026)
von: Xie, Yingsha, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
von: Zhang, Dongcheng, et al.
Veröffentlicht: (2026) -
Reasoning Can Be Restored by Correcting a Few Decision Tokens
von: Shen, Changshuo, et al.
Veröffentlicht: (2026) -
On Reasoning Strength Planning in Large Reasoning Models
von: Sheng, Leheng, et al.
Veröffentlicht: (2025) -
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
von: Zhang, Yi, et al.
Veröffentlicht: (2025) -
Language Representations Can be What Recommenders Need: Findings and Potentials
von: Sheng, Leheng, et al.
Veröffentlicht: (2024)