Watch Before You Answer: Learning from Visually Grounded Post-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuxuan, Hwang, EunJeong, Zhang, Huaisong, Du, Penghui, Jia, Yiming, Jiang, Dongfu, He, Xuan, Zhang, Shenhui, Nie, Ping, West, Peter, Allen, Kelsey R.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915919268151296
author Zhang, Yuxuan
Hwang, EunJeong
Zhang, Huaisong
Du, Penghui
Jia, Yiming
Jiang, Dongfu
He, Xuan
Zhang, Shenhui
Nie, Ping
West, Peter
Allen, Kelsey R.
author_facet Zhang, Yuxuan
Hwang, EunJeong
Zhang, Huaisong
Du, Penghui
Jia, Yiming
Jiang, Dongfu
He, Xuan
Zhang, Shenhui
Nie, Ping
West, Peter
Allen, Kelsey R.
contents It is critical for vision-language models (VLMs) to comprehensively understand visual, temporal, and textual cues. However, despite rapid progress in multimodal modeling, video understanding performance still lags behind text-based reasoning. In this work, we find that progress is even worse than previously assumed: commonly reported long video understanding benchmarks contain 40-60% of questions that can be answered using text cues alone. Furthermore, we find that these issues are also pervasive in widely used post-training datasets, potentially undercutting the ability of post-training to improve VLM video understanding performance. Guided by this observation, we introduce VidGround as a simple yet effective solution: using only the actual visually grounded questions without any linguistic biases for post-training. When used in tandem with RL-based post-training algorithms, this simple technique improves performance by up to 6.2 points relative to using the full dataset, while using only 69.1% of the original post-training data. Moreover, we show that data curation with a simple post-training algorithm outperforms several more complex post-training techniques, highlighting that data quality is a major bottleneck for improving video understanding in VLMs. These results underscore the importance of curating post-training data and evaluation benchmarks that truly require visual grounding to advance the development of more capable VLMs. Project page: http://vidground.etuagi.com.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05117
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Watch Before You Answer: Learning from Visually Grounded Post-Training
Zhang, Yuxuan
Hwang, EunJeong
Zhang, Huaisong
Du, Penghui
Jia, Yiming
Jiang, Dongfu
He, Xuan
Zhang, Shenhui
Nie, Ping
West, Peter
Allen, Kelsey R.
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
It is critical for vision-language models (VLMs) to comprehensively understand visual, temporal, and textual cues. However, despite rapid progress in multimodal modeling, video understanding performance still lags behind text-based reasoning. In this work, we find that progress is even worse than previously assumed: commonly reported long video understanding benchmarks contain 40-60% of questions that can be answered using text cues alone. Furthermore, we find that these issues are also pervasive in widely used post-training datasets, potentially undercutting the ability of post-training to improve VLM video understanding performance. Guided by this observation, we introduce VidGround as a simple yet effective solution: using only the actual visually grounded questions without any linguistic biases for post-training. When used in tandem with RL-based post-training algorithms, this simple technique improves performance by up to 6.2 points relative to using the full dataset, while using only 69.1% of the original post-training data. Moreover, we show that data curation with a simple post-training algorithm outperforms several more complex post-training techniques, highlighting that data quality is a major bottleneck for improving video understanding in VLMs. These results underscore the importance of curating post-training data and evaluation benchmarks that truly require visual grounding to advance the development of more capable VLMs. Project page: http://vidground.etuagi.com.
title Watch Before You Answer: Learning from Visually Grounded Post-Training
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.05117