Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hou, Wenjin, Peng, Shangpin, Wang, Weinong, Ruan, Zheng, Zhang, Yue, Zhou, Zhenglin, Gao, Mingqi, Chen, Yifei, Wang, Kaiqi, Yang, Hongming, Zhang, Chengquan, Tian, Zhuotao, Hu, Han, Yang, Yi, Wu, Fei, Fan, Hehe
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911648512475136
author Hou, Wenjin
Peng, Shangpin
Wang, Weinong
Ruan, Zheng
Zhang, Yue
Zhou, Zhenglin
Gao, Mingqi
Chen, Yifei
Wang, Kaiqi
Yang, Hongming
Zhang, Chengquan
Tian, Zhuotao
Hu, Han
Yang, Yi
Wu, Fei
Fan, Hehe
author_facet Hou, Wenjin
Peng, Shangpin
Wang, Weinong
Ruan, Zheng
Zhang, Yue
Zhou, Zhenglin
Gao, Mingqi
Chen, Yifei
Wang, Kaiqi
Yang, Hongming
Zhang, Chengquan
Tian, Zhuotao
Hu, Han
Yang, Yi
Wu, Fei
Fan, Hehe
contents On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model. Despite its empirical success, the conditions under which OPD yields reliable improvement remain poorly understood. In this work, we identify two fundamental bottlenecks that limit effective OPD: insufficient exploration of informative states and unreliable teacher supervision for student rollouts. Building on this insight, we propose Uni-OPD, a unified OPD framework that generalizes across Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), centered on a dual-perspective optimization strategy. Specifically, from the student's perspective, we adopt two data balancing strategies to promote exploration of informative student-generated states during training. From the teacher's perspective, we show that reliable supervision hinges on whether aggregated token-level guidance remains order-consistent with the outcome reward. To this end, we develop an outcome-guided margin calibration mechanism to restore order consistency between correct and incorrect trajectories. We conduct extensive experiments on 5 domains and 16 benchmarks covering diverse settings, including single-teacher and multi-teacher distillation across LLMs and MLLMs, strong-to-weak distillation, and cross-modal distillation. Our results verify the effectiveness and versatility of Uni-OPD and provide practical insights into reliable OPD.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03677
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Hou, Wenjin
Peng, Shangpin
Wang, Weinong
Ruan, Zheng
Zhang, Yue
Zhou, Zhenglin
Gao, Mingqi
Chen, Yifei
Wang, Kaiqi
Yang, Hongming
Zhang, Chengquan
Tian, Zhuotao
Hu, Han
Yang, Yi
Wu, Fei
Fan, Hehe
Machine Learning
On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model. Despite its empirical success, the conditions under which OPD yields reliable improvement remain poorly understood. In this work, we identify two fundamental bottlenecks that limit effective OPD: insufficient exploration of informative states and unreliable teacher supervision for student rollouts. Building on this insight, we propose Uni-OPD, a unified OPD framework that generalizes across Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), centered on a dual-perspective optimization strategy. Specifically, from the student's perspective, we adopt two data balancing strategies to promote exploration of informative student-generated states during training. From the teacher's perspective, we show that reliable supervision hinges on whether aggregated token-level guidance remains order-consistent with the outcome reward. To this end, we develop an outcome-guided margin calibration mechanism to restore order consistency between correct and incorrect trajectories. We conduct extensive experiments on 5 domains and 16 benchmarks covering diverse settings, including single-teacher and multi-teacher distillation across LLMs and MLLMs, strong-to-weak distillation, and cross-modal distillation. Our results verify the effectiveness and versatility of Uni-OPD and provide practical insights into reliable OPD.
title Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
topic Machine Learning
url https://arxiv.org/abs/2605.03677