Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yaxuan, Zuo, Yuxin, He, Bingxiang, Zhang, Jinqian, Xiao, Chaojun, Qian, Cheng, Yu, Tianyu, Gao, Huan-ang, Yang, Wenkai, Liu, Zhiyuan, Ding, Ning
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914475211227136
author Li, Yaxuan
Zuo, Yuxin
He, Bingxiang
Zhang, Jinqian
Xiao, Chaojun
Qian, Cheng
Yu, Tianyu
Gao, Huan-ang
Yang, Wenkai
Liu, Zhiyuan
Ding, Ning
author_facet Li, Yaxuan
Zuo, Yuxin
He, Bingxiang
Zhang, Jinqian
Xiao, Chaojun
Qian, Cheng
Yu, Tianyu
Gao, Huan-ang
Yang, Wenkai
Liu, Zhiyuan
Ding, Ning
contents On-policy distillation (OPD) has become a core technique in the post-training of large language models, yet its training dynamics remain poorly understood. This paper provides a systematic investigation of OPD dynamics and mechanisms. We first identify that two conditions govern whether OPD succeeds or fails: (i) the student and teacher should share compatible thinking patterns; and (ii) even with consistent thinking patterns and higher scores, the teacher must offer genuinely new capabilities beyond what the student has seen during training. We validate these findings through weak-to-strong reverse distillation, showing that same-family 1.5B and 7B teachers are distributionally indistinguishable from the student's perspective. Probing into the token-level mechanism, we show that successful OPD is characterized by progressive alignment on high-probability tokens at student-visited states, a small shared token set that concentrates most of the probability mass (97%-99%). We further propose two practical strategies to recover failing OPD: off-policy cold start and teacher-aligned prompt selection. Finally, we show that OPD's apparent free lunch of dense token-level reward comes at a cost, raising the question of whether OPD can scale to long-horizon distillation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13016
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Li, Yaxuan
Zuo, Yuxin
He, Bingxiang
Zhang, Jinqian
Xiao, Chaojun
Qian, Cheng
Yu, Tianyu
Gao, Huan-ang
Yang, Wenkai
Liu, Zhiyuan
Ding, Ning
Machine Learning
Artificial Intelligence
Computation and Language
On-policy distillation (OPD) has become a core technique in the post-training of large language models, yet its training dynamics remain poorly understood. This paper provides a systematic investigation of OPD dynamics and mechanisms. We first identify that two conditions govern whether OPD succeeds or fails: (i) the student and teacher should share compatible thinking patterns; and (ii) even with consistent thinking patterns and higher scores, the teacher must offer genuinely new capabilities beyond what the student has seen during training. We validate these findings through weak-to-strong reverse distillation, showing that same-family 1.5B and 7B teachers are distributionally indistinguishable from the student's perspective. Probing into the token-level mechanism, we show that successful OPD is characterized by progressive alignment on high-probability tokens at student-visited states, a small shared token set that concentrates most of the probability mass (97%-99%). We further propose two practical strategies to recover failing OPD: off-policy cold start and teacher-aligned prompt selection. Finally, we show that OPD's apparent free lunch of dense token-level reward comes at a cost, raising the question of whether OPD can scale to long-horizon distillation.
title Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.13016