Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yaxuan, Zuo, Yuxin, He, Bingxiang, Zhang, Jinqian, Xiao, Chaojun, Qian, Cheng, Yu, Tianyu, Gao, Huan-ang, Yang, Wenkai, Liu, Zhiyuan, Ding, Ning |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
by: He, Bingxiang, et al.
Published: (2025)
by: He, Bingxiang, et al.
Published: (2025)
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
by: He, Bingxiang, et al.
Published: (2024)
by: He, Bingxiang, et al.
Published: (2024)
Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
by: Guo, Yiju, et al.
Published: (2025)
by: Guo, Yiju, et al.
Published: (2025)
How Far Can Unsupervised RLVR Scale LLM Training?
by: He, Bingxiang, et al.
Published: (2026)
by: He, Bingxiang, et al.
Published: (2026)
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
by: Hou, Wenjin, et al.
Published: (2026)
by: Hou, Wenjin, et al.
Published: (2026)
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
by: Cheng, Qianjia, et al.
Published: (2026)
by: Cheng, Qianjia, et al.
Published: (2026)
3D Dynamics-Aware Manipulation: Endowing Manipulation Policies with 3D Foresight
by: He, Yuxin, et al.
Published: (2025)
by: He, Yuxin, et al.
Published: (2025)
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
by: Kwon, Soo Min, et al.
Published: (2026)
by: Kwon, Soo Min, et al.
Published: (2026)
Student-in-the-Loop Chain-of-Thought Distillation via Generation-Time Selection
by: He, Chaoqun, et al.
Published: (2026)
by: He, Chaoqun, et al.
Published: (2026)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
by: Gao, Cheng, et al.
Published: (2025)
by: Gao, Cheng, et al.
Published: (2025)
Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate Pairs
by: Gao, Cheng, et al.
Published: (2024)
by: Gao, Cheng, et al.
Published: (2024)
LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
by: Hu, Zhiyuan, et al.
Published: (2024)
by: Hu, Zhiyuan, et al.
Published: (2024)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
by: Zhao, Hanyang, et al.
Published: (2026)
by: Zhao, Hanyang, et al.
Published: (2026)
Diffusion-based Visual Anagram as Multi-task Learning
by: Xu, Zhiyuan, et al.
Published: (2024)
by: Xu, Zhiyuan, et al.
Published: (2024)
Veri-R1: Toward Precise and Faithful Claim Verification via Online Reinforcement Learning
by: He, Qi, et al.
Published: (2025)
by: He, Qi, et al.
Published: (2025)
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
by: Xiao, Yuxin, et al.
Published: (2024)
by: Xiao, Yuxin, et al.
Published: (2024)
The Elephant in the Room: Rethinking the Usage of Pre-trained Language Model in Sequential Recommendation
by: Qu, Zekai, et al.
Published: (2024)
by: Qu, Zekai, et al.
Published: (2024)
HFedCKD: Toward Robust Heterogeneous Federated Learning via Data-free Knowledge Distillation and Two-way Contrast
by: Zheng, Yiting, et al.
Published: (2025)
by: Zheng, Yiting, et al.
Published: (2025)
CubeBench: Diagnosing Interactive, Long-Horizon Spatial Reasoning Under Partial Observations
by: Gao, Huan-ang, et al.
Published: (2025)
by: Gao, Huan-ang, et al.
Published: (2025)
Draft-OPD: On-Policy Distillation for Speculative Draft Models
by: Lei, Haodi, et al.
Published: (2026)
by: Lei, Haodi, et al.
Published: (2026)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
by: Hua, Peichun, et al.
Published: (2025)
by: Hua, Peichun, et al.
Published: (2025)
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
Distilling Rule-based Knowledge into Large Language Models
by: Yang, Wenkai, et al.
Published: (2023)
by: Yang, Wenkai, et al.
Published: (2023)
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Rethinking Post-Training Recipes for Multimodal Time-Series Forecasting
by: Liu, Haoxin, et al.
Published: (2026)
by: Liu, Haoxin, et al.
Published: (2026)
Dual-frame Fluid Motion Estimation with Test-time Optimization and Zero-divergence Loss
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
TPFL: A Trustworthy Personalized Federated Learning Framework via Subjective Logic
by: Chen, Jinqian, et al.
Published: (2024)
by: Chen, Jinqian, et al.
Published: (2024)
TimeRecipe: A Time-Series Forecasting Recipe via Benchmarking Module Level Effectiveness
by: Zhao, Zhiyuan, et al.
Published: (2025)
by: Zhao, Zhiyuan, et al.
Published: (2025)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
by: Ding, Ken
Published: (2026)
by: Ding, Ken
Published: (2026)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
by: Chen, Yingfa, et al.
Published: (2026)
by: Chen, Yingfa, et al.
Published: (2026)
A Study on the Impact of Data Assets on Corporate ESG Performance—— Empirical analysis based on A-share listed companies in China
by: Wang, Yiyu, et al.
Published: (2026)
by: Wang, Yiyu, et al.
Published: (2026)
Toward Suppliers' Green Innovation: The Role of Buyers' Environmental Information Disclosure
by: Bingxiang Li, et al.
Published: (2025)
by: Bingxiang Li, et al.
Published: (2025)
Importance-Weighted Domain Adaptation for Sound Source Tracking
by: Zhong, Bingxiang, et al.
Published: (2025)
by: Zhong, Bingxiang, et al.
Published: (2025)
AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset
by: He, Bingxiang, et al.
Published: (2025)
by: He, Bingxiang, et al.
Published: (2025)
Rethinking Spiking Neural Networks from an Ensemble Learning Perspective
by: Ding, Yongqi, et al.
Published: (2025)
by: Ding, Yongqi, et al.
Published: (2025)
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
by: Nguyen, Thong, et al.
Published: (2025)
by: Nguyen, Thong, et al.
Published: (2025)
Step Rejection Fine-Tuning: A Practical Distillation Recipe
by: Slinko, Igor, et al.
Published: (2026)
by: Slinko, Igor, et al.
Published: (2026)
Similar Items
-
JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
by: He, Bingxiang, et al.
Published: (2025) -
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
by: He, Bingxiang, et al.
Published: (2024) -
Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
by: Guo, Yiju, et al.
Published: (2025) -
How Far Can Unsupervised RLVR Scale LLM Training?
by: He, Bingxiang, et al.
Published: (2026) -
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
by: Hou, Wenjin, et al.
Published: (2026)