VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bousselham, Walid, Kuehne, Hilde, Schmid, Cordelia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026)
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026)
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
von: Bousselham, Walid, et al.
Veröffentlicht: (2026)
von: Bousselham, Walid, et al.
Veröffentlicht: (2026)
LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity
von: Bousselham, Walid, et al.
Veröffentlicht: (2024)
von: Bousselham, Walid, et al.
Veröffentlicht: (2024)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
von: Bousselham, Walid, et al.
Veröffentlicht: (2024)
von: Bousselham, Walid, et al.
Veröffentlicht: (2024)
VideoGEM: Training-free Action Grounding in Videos
von: Vogel, Felix, et al.
Veröffentlicht: (2025)
von: Vogel, Felix, et al.
Veröffentlicht: (2025)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
von: Garcia, Ricardo, et al.
Veröffentlicht: (2024)
von: Garcia, Ricardo, et al.
Veröffentlicht: (2024)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
von: Pacaud, Paul, et al.
Veröffentlicht: (2025)
von: Pacaud, Paul, et al.
Veröffentlicht: (2025)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
von: Chen, Shizhe, et al.
Veröffentlicht: (2026)
von: Chen, Shizhe, et al.
Veröffentlicht: (2026)
Retrieval-Enhanced Contrastive Vision-Text Models
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023)
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023)
TimeLogic: A Temporal Logic Benchmark for Video QA
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
von: Ventura, Lucas, et al.
Veröffentlicht: (2025)
von: Ventura, Lucas, et al.
Veröffentlicht: (2025)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
von: Shvetsova, Nina, et al.
Veröffentlicht: (2023)
von: Shvetsova, Nina, et al.
Veröffentlicht: (2023)
Learning Correlation Structures for Vision Transformers
von: Kim, Manjin, et al.
Veröffentlicht: (2024)
von: Kim, Manjin, et al.
Veröffentlicht: (2024)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
von: Chen, Shizhe, et al.
Veröffentlicht: (2025)
von: Chen, Shizhe, et al.
Veröffentlicht: (2025)
BrickNet: Graph-Backed Generative Brick Assembly
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
TTRV: Test-Time Reinforcement Learning for Vision Language Models
von: Singh, Akshit, et al.
Veröffentlicht: (2025)
von: Singh, Akshit, et al.
Veröffentlicht: (2025)
Learning text-to-video retrieval from image captioning
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
von: Ventura, Lucas, et al.
Veröffentlicht: (2024)
Continual Learning in Vision-Language Models via Aligned Model Merging
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
Visual-Advantage On-Policy Distillation for Vision-Language Models
von: Liu, Ruiqi, et al.
Veröffentlicht: (2026)
von: Liu, Ruiqi, et al.
Veröffentlicht: (2026)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
von: Araujo, Edson, et al.
Veröffentlicht: (2026)
von: Araujo, Edson, et al.
Veröffentlicht: (2026)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2024)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2025)
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2025)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
Dual Guidance Semi-Supervised Action Detection
von: Singh, Ankit, et al.
Veröffentlicht: (2025)
von: Singh, Ankit, et al.
Veröffentlicht: (2025)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
WPT: World-to-Policy Transfer via Online World Model Distillation
von: Jiang, Guangfeng, et al.
Veröffentlicht: (2025)
von: Jiang, Guangfeng, et al.
Veröffentlicht: (2025)
Dense Video Object Captioning from Disjoint Supervision
von: Zhou, Xingyi, et al.
Veröffentlicht: (2023)
von: Zhou, Xingyi, et al.
Veröffentlicht: (2023)
InteractVLM: 3D Interaction Reasoning from 2D Foundational Models
von: Dwivedi, Sai Kumar, et al.
Veröffentlicht: (2025)
von: Dwivedi, Sai Kumar, et al.
Veröffentlicht: (2025)
Dense Optical Tracking: Connecting the Dots
von: Moing, Guillaume Le, et al.
Veröffentlicht: (2023)
von: Moing, Guillaume Le, et al.
Veröffentlicht: (2023)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
von: Hu, Yushi, et al.
Veröffentlicht: (2023)
von: Hu, Yushi, et al.
Veröffentlicht: (2023)
MetricNet: Recovering Metric Scale in Generative Navigation Policies
von: Nayak, Abhijeet, et al.
Veröffentlicht: (2025)
von: Nayak, Abhijeet, et al.
Veröffentlicht: (2025)
WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition
von: Bock, Marius, et al.
Veröffentlicht: (2023)
von: Bock, Marius, et al.
Veröffentlicht: (2023)
DAIT: Distillation from Vision-Language Models to Lightweight Classifiers with Adaptive Intermediate Teacher Transfer
von: He, Zhengxu, et al.
Veröffentlicht: (2026)
von: He, Zhengxu, et al.
Veröffentlicht: (2026)
SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces
von: Wu, Guande, et al.
Veröffentlicht: (2025)
von: Wu, Guande, et al.
Veröffentlicht: (2025)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
SUGAR: Pre-training 3D Visual Representations for Robotics
von: Chen, Shizhe, et al.
Veröffentlicht: (2024)
von: Chen, Shizhe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
von: Jahagirdar, Soumya, et al.
Veröffentlicht: (2026) -
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
von: Bousselham, Walid, et al.
Veröffentlicht: (2026) -
LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity
von: Bousselham, Walid, et al.
Veröffentlicht: (2024) -
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025) -
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
von: Bousselham, Walid, et al.
Veröffentlicht: (2024)