Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Dongcheng, Zhang, Yi, Chen, Yuxin, Zhang, An, Wang, Xiang, Lu, Chaochao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Internalizing Safety Understanding in Large Reasoning Models via Verification
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
von: Zhoubian, Sining, et al.
Veröffentlicht: (2025)
von: Zhoubian, Sining, et al.
Veröffentlicht: (2025)
EEG-ReMinD: Enhancing Neurodegenerative EEG Decoding through Self-Supervised State Reconstruction-Primed Riemannian Dynamics
von: Wang, Zirui, et al.
Veröffentlicht: (2025)
von: Wang, Zirui, et al.
Veröffentlicht: (2025)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
Self-Evolving Curriculum for LLM Reasoning
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2025)
Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networks
von: Wang, Jihang, et al.
Veröffentlicht: (2025)
von: Wang, Jihang, et al.
Veröffentlicht: (2025)
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
von: Wang, Yuanfu, et al.
Veröffentlicht: (2026)
von: Wang, Yuanfu, et al.
Veröffentlicht: (2026)
Learning Robust Reasoning through Guided Adversarial Self-Play
von: Li, Shuozhe, et al.
Veröffentlicht: (2026)
von: Li, Shuozhe, et al.
Veröffentlicht: (2026)
CoRe-ECG: Advancing Self-Supervised Representation Learning for 12-Lead ECG via Contrastive and Reconstructive Synergy
von: Qin, Zehao, et al.
Veröffentlicht: (2026)
von: Qin, Zehao, et al.
Veröffentlicht: (2026)
SafeCoT: Improving VLM Safety with Minimal Reasoning
von: Ma, Jiachen, et al.
Veröffentlicht: (2025)
von: Ma, Jiachen, et al.
Veröffentlicht: (2025)
Dynamic Adversarial Reinforcement Learning for Robust Multimodal Large Language Models
von: Bao, Yicheng, et al.
Veröffentlicht: (2026)
von: Bao, Yicheng, et al.
Veröffentlicht: (2026)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
Safe-Support Q-Learning: Learning without Unsafe Exploration
von: Lim, Yeeun, et al.
Veröffentlicht: (2026)
von: Lim, Yeeun, et al.
Veröffentlicht: (2026)
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
von: Shen, Sicheng, et al.
Veröffentlicht: (2026)
von: Shen, Sicheng, et al.
Veröffentlicht: (2026)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
Generalizable Multimodal Large Language Model Editing via Invariant Trajectory Learning
von: Su, Jiajie, et al.
Veröffentlicht: (2026)
von: Su, Jiajie, et al.
Veröffentlicht: (2026)
SelfBC: Self Behavior Cloning for Offline Reinforcement Learning
von: Liu, Shirong, et al.
Veröffentlicht: (2024)
von: Liu, Shirong, et al.
Veröffentlicht: (2024)
Learning Diverse Policies with Soft Self-Generated Guidance
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
Self-Supervised Learning for Time Series: Contrastive or Generative?
von: Liu, Ziyu, et al.
Veröffentlicht: (2024)
von: Liu, Ziyu, et al.
Veröffentlicht: (2024)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
Context-Enhanced Multi-View Trajectory Representation Learning: Bridging the Gap through Self-Supervised Models
von: Qian, Tangwen, et al.
Veröffentlicht: (2024)
von: Qian, Tangwen, et al.
Veröffentlicht: (2024)
RePO: Replay-Enhanced Policy Optimization
von: Li, Siheng, et al.
Veröffentlicht: (2025)
von: Li, Siheng, et al.
Veröffentlicht: (2025)
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
von: Zhou, Huilin, et al.
Veröffentlicht: (2026)
von: Zhou, Huilin, et al.
Veröffentlicht: (2026)
Meta-Cognitive Reinforcement Learning with Self-Doubt and Recovery
von: Zhang, Zhipeng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhipeng, et al.
Veröffentlicht: (2026)
Self-rewarding correction for mathematical reasoning
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
RL-STaR: Theoretical Analysis of Reinforcement Learning Frameworks for Self-Taught Reasoner
von: Chang, Fu-Chieh, et al.
Veröffentlicht: (2024)
von: Chang, Fu-Chieh, et al.
Veröffentlicht: (2024)
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
Self-Improved Learning for Scalable Neural Combinatorial Optimization
von: Luo, Fu, et al.
Veröffentlicht: (2024)
von: Luo, Fu, et al.
Veröffentlicht: (2024)
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
Self-Controlled Dynamic Expansion Model for Continual Learning
von: Wu, Runqing, et al.
Veröffentlicht: (2025)
von: Wu, Runqing, et al.
Veröffentlicht: (2025)
FedEmb: A Vertical and Hybrid Federated Learning Algorithm using Network And Feature Embedding Aggregation
von: Meng, Fanfei, et al.
Veröffentlicht: (2023)
von: Meng, Fanfei, et al.
Veröffentlicht: (2023)
Learning from Partial Chain-of-Thought via Truncated-Reasoning Self-Distillation
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2026)
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2026)
Prototypical Self-Explainable Models Without Re-training
von: Gautam, Srishti, et al.
Veröffentlicht: (2023)
von: Gautam, Srishti, et al.
Veröffentlicht: (2023)
Adaptive Self-supervised Robust Clustering for Unstructured Data with Unknown Cluster Number
von: Ding, Chen-Lu, et al.
Veröffentlicht: (2024)
von: Ding, Chen-Lu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Internalizing Safety Understanding in Large Reasoning Models via Verification
von: Zhang, Yi, et al.
Veröffentlicht: (2026) -
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
von: Shen, Guobin, et al.
Veröffentlicht: (2026) -
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
von: Zhoubian, Sining, et al.
Veröffentlicht: (2025) -
EEG-ReMinD: Enhancing Neurodegenerative EEG Decoding through Self-Supervised State Reconstruction-Primed Riemannian Dynamics
von: Wang, Zirui, et al.
Veröffentlicht: (2025) -
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)