VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bousselham, Walid, Kuehne, Hilde, Schmid, Cordelia
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914117558730752
author Bousselham, Walid
Kuehne, Hilde
Schmid, Cordelia
author_facet Bousselham, Walid
Kuehne, Hilde
Schmid, Cordelia
contents Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoning data. Conversely, text-based reasoning resources are abundant and scalable, but it is still an open question how to leveraging them for VLM reasoning. To address this problem, we propose VOLD, a framework to transfer reasoning capabilities from text-only teacher models to VLM student models. To this end, VOLD combines reinforcement learning via Group Relative Policy Optimization (GRPO) with on-policy distillation, which allows the student reasoning traces to be guided by the teacher model, resulting in a significant gain over using GRPO alone. We further show that a cold-start alignment is essential for an effective transfer during the online training phase in this scenario and that without sufficient distributional alignment between teacher and student, on-policy distillation fails to provide meaningful guidance. We evaluate VOLD across diverse benchmarks including MMMU-Pro, MathVision, MathVista, and LogicVista, showing that VOLD outperforms the baseline model significantly and improves over the state of the art by a margin. Our ablation shows the importance of a cold-start alignment via SFT for on-policy distillation with a text-only teacher.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23497
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
Bousselham, Walid
Kuehne, Hilde
Schmid, Cordelia
Computer Vision and Pattern Recognition
Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoning data. Conversely, text-based reasoning resources are abundant and scalable, but it is still an open question how to leveraging them for VLM reasoning. To address this problem, we propose VOLD, a framework to transfer reasoning capabilities from text-only teacher models to VLM student models. To this end, VOLD combines reinforcement learning via Group Relative Policy Optimization (GRPO) with on-policy distillation, which allows the student reasoning traces to be guided by the teacher model, resulting in a significant gain over using GRPO alone. We further show that a cold-start alignment is essential for an effective transfer during the online training phase in this scenario and that without sufficient distributional alignment between teacher and student, on-policy distillation fails to provide meaningful guidance. We evaluate VOLD across diverse benchmarks including MMMU-Pro, MathVision, MathVista, and LogicVista, showing that VOLD outperforms the baseline model significantly and improves over the state of the art by a margin. Our ablation shows the importance of a cold-start alignment via SFT for on-policy distillation with a text-only teacher.
title VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.23497