Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Xin, AI, Cliveb, Yang, Kai, Chen, Tianhao, Wang, Yang, Yang, Saiyong, Yang, Can |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization
von: He, Junlin, et al.
Veröffentlicht: (2026)
von: He, Junlin, et al.
Veröffentlicht: (2026)
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026)
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026)
Think Beyond Size: Adaptive Prompting for More Effective Reasoning
von: R, Kamesh
Veröffentlicht: (2024)
von: R, Kamesh
Veröffentlicht: (2024)
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
von: Ko, Jongwoo, et al.
Veröffentlicht: (2026)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2026)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
von: Wang, Xiao
Veröffentlicht: (2026)
von: Wang, Xiao
Veröffentlicht: (2026)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
von: Zhao, Siyan, et al.
Veröffentlicht: (2026)
von: Zhao, Siyan, et al.
Veröffentlicht: (2026)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
von: Xu, Xin, et al.
Veröffentlicht: (2026)
von: Xu, Xin, et al.
Veröffentlicht: (2026)
Reverse Thinking Makes LLMs Stronger Reasoners
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model
von: Tang, Chenming, et al.
Veröffentlicht: (2026)
von: Tang, Chenming, et al.
Veröffentlicht: (2026)
Efficient Reasoning with Balanced Thinking
von: Li, Yulin, et al.
Veröffentlicht: (2026)
von: Li, Yulin, et al.
Veröffentlicht: (2026)
Effectively Controlling Reasoning Models through Thinking Intervention
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
Efficient Reasoning with Hidden Thinking
von: Shen, Xuan, et al.
Veröffentlicht: (2025)
von: Shen, Xuan, et al.
Veröffentlicht: (2025)
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
von: Yang, Wang, et al.
Veröffentlicht: (2026)
von: Yang, Wang, et al.
Veröffentlicht: (2026)
Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
Relative Kinetic Utility for Reasoning-Aware Structural Pruning in Large Language Models
von: Qian, Tianhao
Veröffentlicht: (2026)
von: Qian, Tianhao
Veröffentlicht: (2026)
Think Outside the Policy: In-Context Steered Policy Optimization
von: Huang, Hsiu-Yuan, et al.
Veröffentlicht: (2025)
von: Huang, Hsiu-Yuan, et al.
Veröffentlicht: (2025)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
AdapThink: Adaptive Thinking Preferences for Reasoning Language Model
von: Wan, Xu, et al.
Veröffentlicht: (2025)
von: Wan, Xu, et al.
Veröffentlicht: (2025)
Are Reasoning Models More Prone to Hallucination?
von: Yao, Zijun, et al.
Veröffentlicht: (2025)
von: Yao, Zijun, et al.
Veröffentlicht: (2025)
Reasoning Models Don't Just Think Longer, They Move Differently
von: Gjølbye, Anders, et al.
Veröffentlicht: (2026)
von: Gjølbye, Anders, et al.
Veröffentlicht: (2026)
Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
von: Yakushev, George, et al.
Veröffentlicht: (2025)
von: Yakushev, George, et al.
Veröffentlicht: (2025)
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
von: Xu, Xin, et al.
Veröffentlicht: (2026)
von: Xu, Xin, et al.
Veröffentlicht: (2026)
What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
Agentic-R1: Distilled Dual-Strategy Reasoning
von: Du, Weihua, et al.
Veröffentlicht: (2025)
von: Du, Weihua, et al.
Veröffentlicht: (2025)
An Analysis for Reasoning Bias of Language Models with Small Initialization
von: Yao, Junjie, et al.
Veröffentlicht: (2025)
von: Yao, Junjie, et al.
Veröffentlicht: (2025)
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
Structural Rationale Distillation via Reasoning Space Compression
von: Yang, Jialin, et al.
Veröffentlicht: (2026)
von: Yang, Jialin, et al.
Veröffentlicht: (2026)
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models
von: Wu, Da, et al.
Veröffentlicht: (2023)
von: Wu, Da, et al.
Veröffentlicht: (2023)
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
Recursive Models for Long-Horizon Reasoning
von: Yang, Chenxiao, et al.
Veröffentlicht: (2026)
von: Yang, Chenxiao, et al.
Veröffentlicht: (2026)
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
von: Luo, Haipeng, et al.
Veröffentlicht: (2025)
von: Luo, Haipeng, et al.
Veröffentlicht: (2025)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
von: Qu, Yun, et al.
Veröffentlicht: (2026)
von: Qu, Yun, et al.
Veröffentlicht: (2026)
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner
von: He, Tao, et al.
Veröffentlicht: (2025)
von: He, Tao, et al.
Veröffentlicht: (2025)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
von: Yang, Wenkai, et al.
Veröffentlicht: (2026) -
Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization
von: He, Junlin, et al.
Veröffentlicht: (2026) -
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026) -
Think Beyond Size: Adaptive Prompting for More Effective Reasoning
von: R, Kamesh
Veröffentlicht: (2024) -
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)