Visual-Advantage On-Policy Distillation for Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Ruiqi, Lv, Xiaolei, Li, Gengsheng, Zhu, Ximo, Wang, Zhiheng, Zhang, Zhengbo, Chen, Junkai, Li, Zhiheng, Li, Bo, Gao, Jun, Wu, Shu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910243302146048
author Liu, Ruiqi
Lv, Xiaolei
Li, Gengsheng
Zhu, Ximo
Wang, Zhiheng
Zhang, Zhengbo
Chen, Junkai
Li, Zhiheng
Li, Bo
Gao, Jun
Wu, Shu
author_facet Liu, Ruiqi
Lv, Xiaolei
Li, Gengsheng
Zhu, Ximo
Wang, Zhiheng
Zhang, Zhengbo
Chen, Junkai
Li, Zhiheng
Li, Bo
Gao, Jun
Wu, Shu
contents On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-policy distillation can improve a student's output quality while failing to strengthen its reliance on visual input: on vision-critical tokens, the student's predictions remain largely unchanged whether or not fine-grained visual detail is present, even though the teacher's predictions depend heavily on it.To make this difference observable, we introduce visual advantage (VA), the token-level log-probability difference when the teacher scores a student-generated rollout with versus without access to fine-grained visual detail. VA is concentrated in a small minority of tokens, and these high-VA tokens are the ones that actually carry the visual supervision signal. This motivates a distillation objective that treats them differently from language scaffolding, so their contribution is not diluted by the abundant surrounding language tokens.We propose Visual-Advantage On-Policy Distillation (VA-OPD), which uses VA at two granularities: rollout-level reweighting by trajectory-averaged VA, and token-level KL averaged within high-VA and low-VA groups separately. We train on two math datasets (Geometry3K and ViRL39K) and evaluate on eight benchmarks covering both mathematical reasoning and visual understanding, across three teacher sizes (4B, 8B, and 32B) on the Qwen3-VL family. VA-OPD improves over standard on-policy distillation on every benchmark, with the gain growing monotonically along both the teacher-size and data-scale axes, suggesting that these factors compound consistently.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21924
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Visual-Advantage On-Policy Distillation for Vision-Language Models
Liu, Ruiqi
Lv, Xiaolei
Li, Gengsheng
Zhu, Ximo
Wang, Zhiheng
Zhang, Zhengbo
Chen, Junkai
Li, Zhiheng
Li, Bo
Gao, Jun
Wu, Shu
Computer Vision and Pattern Recognition
On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-policy distillation can improve a student's output quality while failing to strengthen its reliance on visual input: on vision-critical tokens, the student's predictions remain largely unchanged whether or not fine-grained visual detail is present, even though the teacher's predictions depend heavily on it.To make this difference observable, we introduce visual advantage (VA), the token-level log-probability difference when the teacher scores a student-generated rollout with versus without access to fine-grained visual detail. VA is concentrated in a small minority of tokens, and these high-VA tokens are the ones that actually carry the visual supervision signal. This motivates a distillation objective that treats them differently from language scaffolding, so their contribution is not diluted by the abundant surrounding language tokens.We propose Visual-Advantage On-Policy Distillation (VA-OPD), which uses VA at two granularities: rollout-level reweighting by trajectory-averaged VA, and token-level KL averaged within high-VA and low-VA groups separately. We train on two math datasets (Geometry3K and ViRL39K) and evaluate on eight benchmarks covering both mathematical reasoning and visual understanding, across three teacher sizes (4B, 8B, and 32B) on the Qwen3-VL family. VA-OPD improves over standard on-policy distillation on every benchmark, with the gain growing monotonically along both the teacher-size and data-scale axes, suggesting that these factors compound consistently.
title Visual-Advantage On-Policy Distillation for Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.21924