Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Rang, Miao, Bi, Zhenni, Zhou, Hang, Han, Kai, Wang, Xuechun, Xiao, An, Chen, Xinghao, Wang, Yunhe, Chen, Hanting
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913098464493568
author Rang, Miao
Bi, Zhenni
Zhou, Hang
Han, Kai
Wang, Xuechun
Xiao, An
Chen, Xinghao
Wang, Yunhe
Chen, Hanting
author_facet Rang, Miao
Bi, Zhenni
Zhou, Hang
Han, Kai
Wang, Xuechun
Xiao, An
Chen, Xinghao
Wang, Yunhe
Chen, Hanting
contents Standard knowledge distillation for autoregressive models often suffers from distribution mismatch. While on-policy methods mitigate this by leveraging student-generated outputs, they rely on computationally expensive Reinforcement Learning (RL) frameworks. To improve efficiency, we propose Near-Policy Distillation (NPD), an asynchronous approach that decouples student generation from training. This reformulation enables Supervised Fine-Tuning (SFT) with sequence packing. However, asynchronous updates inevitably introduce policy lag and sample noise, which can cause the behavior to drift from near-policy toward off-policy. To counteract this without sacrificing efficiency, NPD integrates sparse student updates and the $Δ$-IFD filtering mechanism, a heuristic sample selection mechanism that empirically stabilizes the optimization trajectory. By filtering extreme out-of-distribution samples, $Δ$-IFD prevents noise from dominating the gradients, ensuring updates remain within a safe proximal learning zone. Empirically, the NPD framework achieves a 8.1x speedup over on-policy baselines and outperforms SFT by 8.09%. Crucially, by effectively narrowing the exploration space for subsequent RL, our method enables openPangu-Embedded-1B to reach a state-of-the-art score of 68.73%, outperforming the substantially larger Qwen3-1.7B. Codes will be released soon.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05940
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing
Rang, Miao
Bi, Zhenni
Zhou, Hang
Han, Kai
Wang, Xuechun
Xiao, An
Chen, Xinghao
Wang, Yunhe
Chen, Hanting
Machine Learning
Computation and Language
Standard knowledge distillation for autoregressive models often suffers from distribution mismatch. While on-policy methods mitigate this by leveraging student-generated outputs, they rely on computationally expensive Reinforcement Learning (RL) frameworks. To improve efficiency, we propose Near-Policy Distillation (NPD), an asynchronous approach that decouples student generation from training. This reformulation enables Supervised Fine-Tuning (SFT) with sequence packing. However, asynchronous updates inevitably introduce policy lag and sample noise, which can cause the behavior to drift from near-policy toward off-policy. To counteract this without sacrificing efficiency, NPD integrates sparse student updates and the $Δ$-IFD filtering mechanism, a heuristic sample selection mechanism that empirically stabilizes the optimization trajectory. By filtering extreme out-of-distribution samples, $Δ$-IFD prevents noise from dominating the gradients, ensuring updates remain within a safe proximal learning zone. Empirically, the NPD framework achieves a 8.1x speedup over on-policy baselines and outperforms SFT by 8.09%. Crucially, by effectively narrowing the exploration space for subsequent RL, our method enables openPangu-Embedded-1B to reach a state-of-the-art score of 68.73%, outperforming the substantially larger Qwen3-1.7B. Codes will be released soon.
title Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.05940