Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Xiaomin, Hou, Jianheng, Deng, Zheyuan, Zhang, Zhiwei, Li, Taoran, Lu, Binghang, Hu, Bing, Zhao, Yunhan, Hao, Yuexing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning
por: Lu, Binghang, et al.
Publicado: (2026)
por: Lu, Binghang, et al.
Publicado: (2026)
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
por: Tian, Yuan, et al.
Publicado: (2026)
por: Tian, Yuan, et al.
Publicado: (2026)
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
por: Li, Xiaomin, et al.
Publicado: (2025)
por: Li, Xiaomin, et al.
Publicado: (2025)
Dual Identities, Singular Journeys: How Working Tourists' Role Perception Shapes Place Attachment Through Memorable Tourism Experience
por: Yelin Fang, et al.
Publicado: (2026)
por: Yelin Fang, et al.
Publicado: (2026)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
por: Liu, Zheyuan, et al.
Publicado: (2025)
por: Liu, Zheyuan, et al.
Publicado: (2025)
FAAGC: Feature Augmentation on Adaptive Geodesic Curve Based on the shape space theory
por: Han, Yuexing, et al.
Publicado: (2025)
por: Han, Yuexing, et al.
Publicado: (2025)
AdamFLIP: Adaptive Momentum Feedback Linearization Optimization for Hard Constrained PINN Training
por: Lu, Binghang, et al.
Publicado: (2026)
por: Lu, Binghang, et al.
Publicado: (2026)
Dynamics of Apparent Horizon and a Null Comparison Principle
por: An, Xinliang, et al.
Publicado: (2023)
por: An, Xinliang, et al.
Publicado: (2023)
ProbeLLM: Automating Principled Diagnosis of LLM Failures
por: Huang, Yue, et al.
Publicado: (2026)
por: Huang, Yue, et al.
Publicado: (2026)
Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
por: Park, Jonghyun, et al.
Publicado: (2025)
por: Park, Jonghyun, et al.
Publicado: (2025)
Mitigating Overthinking in Large Reasoning Models via Manifold Steering
por: Huang, Yao, et al.
Publicado: (2025)
por: Huang, Yao, et al.
Publicado: (2025)
Data-adaptive Safety Rules for Training Reward Models
por: Li, Xiaomin, et al.
Publicado: (2025)
por: Li, Xiaomin, et al.
Publicado: (2025)
Chain of Mindset: Reasoning with Adaptive Cognitive Modes
por: Jiang, Tianyi, et al.
Publicado: (2026)
por: Jiang, Tianyi, et al.
Publicado: (2026)
Temperature and wind characteristics of Lenghu site for ventilation and structural design of large telescope enclosure
por: Li, Taoran, et al.
Publicado: (2025)
por: Li, Taoran, et al.
Publicado: (2025)
Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
por: Xie, Yingsha, et al.
Publicado: (2026)
por: Xie, Yingsha, et al.
Publicado: (2026)
Mitigating Overthinking in Large Reasoning Language Models via Reasoning Path Deviation Monitoring
por: Guan, Weixin, et al.
Publicado: (2026)
por: Guan, Weixin, et al.
Publicado: (2026)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
por: Zou, Zhengtao, et al.
Publicado: (2025)
por: Zou, Zhengtao, et al.
Publicado: (2025)
Dot1l Regulates the Spontaneous Bone Regeneration of Periosteum‐Derived Stem Cells by Regulating Chac1 Expression
por: Taoran Jiang, et al.
Publicado: (2025)
por: Taoran Jiang, et al.
Publicado: (2025)
fPINN-DeepONet: A Physics-Informed Operator Learning Framework for Multi-term Time-fractional Mixed Diffusion-wave Equations
por: Lu, Binghang, et al.
Publicado: (2026)
por: Lu, Binghang, et al.
Publicado: (2026)
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model
por: Lou, Xinyue, et al.
Publicado: (2025)
por: Lou, Xinyue, et al.
Publicado: (2025)
Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models
por: Liu, Yi, et al.
Publicado: (2025)
por: Liu, Yi, et al.
Publicado: (2025)
Muon with Spectral Guidance: Efficient Optimization for Scientific Machine Learning
por: Lu, Binghang, et al.
Publicado: (2026)
por: Lu, Binghang, et al.
Publicado: (2026)
iPINNER: An Iterative Physics-Informed Neural Network with Ensemble Kalman Filter
por: Lu, Binghang, et al.
Publicado: (2025)
por: Lu, Binghang, et al.
Publicado: (2025)
Morephy-Net: An Evolutionary Multi-objective Optimization for Replica-Exchange-based Physics-informed Neural Operator Learning Networks
por: Lu, Binghang, et al.
Publicado: (2025)
por: Lu, Binghang, et al.
Publicado: (2025)
Neural-POD: A Plug-and-Play Neural Operator Framework for Infinite-Dimensional Functional Nonlinear Proper Orthogonal Decomposition
por: Mou, Changhong, et al.
Publicado: (2026)
por: Mou, Changhong, et al.
Publicado: (2026)
Investigating the Effectiveness of a Socratic Chain-of-Thoughts Reasoning Method for Task Planning in Robotics, A Case Study
por: Bot, Veronica, et al.
Publicado: (2025)
por: Bot, Veronica, et al.
Publicado: (2025)
Understanding mechanisms underlying solar cycle predictability with a general framework
por: Li, Binghang, et al.
Publicado: (2026)
por: Li, Binghang, et al.
Publicado: (2026)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
por: Zhu, Zihao, et al.
Publicado: (2025)
por: Zhu, Zihao, et al.
Publicado: (2025)
Mitigating Cognitive Inertia in Large Reasoning Models via Latent Spike Steering
por: Lee, Seojin, et al.
Publicado: (2026)
por: Lee, Seojin, et al.
Publicado: (2026)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
por: Wang, Zihu, et al.
Publicado: (2025)
por: Wang, Zihu, et al.
Publicado: (2025)
STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
por: Wu, Di, et al.
Publicado: (2026)
por: Wu, Di, et al.
Publicado: (2026)
FAGC:Feature Augmentation on Geodesic Curve in the Pre-Shape Space
por: Han, Yuexing, et al.
Publicado: (2023)
por: Han, Yuexing, et al.
Publicado: (2023)
Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
por: McGregor, Sean, et al.
Publicado: (2025)
por: McGregor, Sean, et al.
Publicado: (2025)
SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering
por: Cao, Zouying, et al.
Publicado: (2024)
por: Cao, Zouying, et al.
Publicado: (2024)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
por: Li, Zihao, et al.
Publicado: (2025)
por: Li, Zihao, et al.
Publicado: (2025)
Advanced Electrolyte Solution for Aqueous Lithium‐Ion Batteries: Extending Their Lifespan
por: Binghang Liu, et al.
Publicado: (2025)
por: Binghang Liu, et al.
Publicado: (2025)
Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs
por: Zhao, Weixiang, et al.
Publicado: (2025)
por: Zhao, Weixiang, et al.
Publicado: (2025)
FAGStyle: Feature Augmentation on Geodesic Surface for Zero-shot Text-guided Diffusion Image Style Transfer
por: Han, Yuexing, et al.
Publicado: (2024)
por: Han, Yuexing, et al.
Publicado: (2024)
Few-shot Image Generation via Information Transfer from the Built Geodesic Surface
por: Han, Yuexing, et al.
Publicado: (2024)
por: Han, Yuexing, et al.
Publicado: (2024)
Lifting the Veil on Composition, Risks, and Mitigations of the Large Language Model Supply Chain
por: Huang, Kaifeng, et al.
Publicado: (2024)
por: Huang, Kaifeng, et al.
Publicado: (2024)
Ejemplares similares
-
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning
por: Lu, Binghang, et al.
Publicado: (2026) -
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
por: Tian, Yuan, et al.
Publicado: (2026) -
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
por: Li, Xiaomin, et al.
Publicado: (2025) -
Dual Identities, Singular Journeys: How Working Tourists' Role Perception Shapes Place Attachment Through Memorable Tourism Experience
por: Yelin Fang, et al.
Publicado: (2026) -
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
por: Liu, Zheyuan, et al.
Publicado: (2025)