Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Jiaze, Yin, Hao, Xu, Haoran, Xu, Boshen, Tan, Wenhui, He, Zewen, Ju, Jianzhong, Luo, Zhenbo, Luan, Jian
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910220369788928
author Li, Jiaze
Yin, Hao
Xu, Haoran
Xu, Boshen
Tan, Wenhui
He, Zewen
Ju, Jianzhong
Luo, Zhenbo
Luan, Jian
author_facet Li, Jiaze
Yin, Hao
Xu, Haoran
Xu, Boshen
Tan, Wenhui
He, Zewen
Ju, Jianzhong
Luo, Zhenbo
Luan, Jian
contents Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet existing GRPO-based methods remain fundamentally constrained by sparse reward signals and substantial computational overhead. We propose Video-OPD, an efficient post-training framework for TVG inspired by recent advances in on-policy distillation. Video-OPD optimizes trajectories sampled directly from the current policy, thereby preserving alignment between training and inference distributions, while a frontier teacher supplies dense, token-level supervision via a reverse KL divergence objective. This formulation preserves the on-policy property critical for mitigating distributional shift, while converting sparse, episode-level feedback into fine-grained, step-wise learning signals. Building on Video-OPD, we introduce Teacher-Validated Disagreement Focusing (TVDF), a lightweight training curriculum that iteratively prioritizes trajectories that are both teacher-reliable and maximally informative for the student, thereby improving training efficiency. Empirical results demonstrate that Video-OPD consistently outperforms GRPO while achieving substantially faster convergence and lower computational cost, establishing on-policy distillation as an effective alternative to conventional reinforcement learning for TVG.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02994
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
Li, Jiaze
Yin, Hao
Xu, Haoran
Xu, Boshen
Tan, Wenhui
He, Zewen
Ju, Jianzhong
Luo, Zhenbo
Luan, Jian
Computer Vision and Pattern Recognition
Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet existing GRPO-based methods remain fundamentally constrained by sparse reward signals and substantial computational overhead. We propose Video-OPD, an efficient post-training framework for TVG inspired by recent advances in on-policy distillation. Video-OPD optimizes trajectories sampled directly from the current policy, thereby preserving alignment between training and inference distributions, while a frontier teacher supplies dense, token-level supervision via a reverse KL divergence objective. This formulation preserves the on-policy property critical for mitigating distributional shift, while converting sparse, episode-level feedback into fine-grained, step-wise learning signals. Building on Video-OPD, we introduce Teacher-Validated Disagreement Focusing (TVDF), a lightweight training curriculum that iteratively prioritizes trajectories that are both teacher-reliable and maximally informative for the student, thereby improving training efficiency. Empirical results demonstrate that Video-OPD consistently outperforms GRPO while achieving substantially faster convergence and lower computational cost, establishing on-policy distillation as an effective alternative to conventional reinforcement learning for TVG.
title Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.02994