Watch Before You Answer: Learning from Visually Grounded Post-Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915919268151296 |
|---|---|
| author | Zhang, Yuxuan Hwang, EunJeong Zhang, Huaisong Du, Penghui Jia, Yiming Jiang, Dongfu He, Xuan Zhang, Shenhui Nie, Ping West, Peter Allen, Kelsey R. |
| author_facet | Zhang, Yuxuan Hwang, EunJeong Zhang, Huaisong Du, Penghui Jia, Yiming Jiang, Dongfu He, Xuan Zhang, Shenhui Nie, Ping West, Peter Allen, Kelsey R. |
| contents | It is critical for vision-language models (VLMs) to comprehensively understand visual, temporal, and textual cues. However, despite rapid progress in multimodal modeling, video understanding performance still lags behind text-based reasoning. In this work, we find that progress is even worse than previously assumed: commonly reported long video understanding benchmarks contain 40-60% of questions that can be answered using text cues alone. Furthermore, we find that these issues are also pervasive in widely used post-training datasets, potentially undercutting the ability of post-training to improve VLM video understanding performance. Guided by this observation, we introduce VidGround as a simple yet effective solution: using only the actual visually grounded questions without any linguistic biases for post-training. When used in tandem with RL-based post-training algorithms, this simple technique improves performance by up to 6.2 points relative to using the full dataset, while using only 69.1% of the original post-training data. Moreover, we show that data curation with a simple post-training algorithm outperforms several more complex post-training techniques, highlighting that data quality is a major bottleneck for improving video understanding in VLMs. These results underscore the importance of curating post-training data and evaluation benchmarks that truly require visual grounding to advance the development of more capable VLMs. Project page: http://vidground.etuagi.com. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_05117 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Watch Before You Answer: Learning from Visually Grounded Post-Training Zhang, Yuxuan Hwang, EunJeong Zhang, Huaisong Du, Penghui Jia, Yiming Jiang, Dongfu He, Xuan Zhang, Shenhui Nie, Ping West, Peter Allen, Kelsey R. Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language It is critical for vision-language models (VLMs) to comprehensively understand visual, temporal, and textual cues. However, despite rapid progress in multimodal modeling, video understanding performance still lags behind text-based reasoning. In this work, we find that progress is even worse than previously assumed: commonly reported long video understanding benchmarks contain 40-60% of questions that can be answered using text cues alone. Furthermore, we find that these issues are also pervasive in widely used post-training datasets, potentially undercutting the ability of post-training to improve VLM video understanding performance. Guided by this observation, we introduce VidGround as a simple yet effective solution: using only the actual visually grounded questions without any linguistic biases for post-training. When used in tandem with RL-based post-training algorithms, this simple technique improves performance by up to 6.2 points relative to using the full dataset, while using only 69.1% of the original post-training data. Moreover, we show that data curation with a simple post-training algorithm outperforms several more complex post-training techniques, highlighting that data quality is a major bottleneck for improving video understanding in VLMs. These results underscore the importance of curating post-training data and evaluation benchmarks that truly require visual grounding to advance the development of more capable VLMs. Project page: http://vidground.etuagi.com. |
| title | Watch Before You Answer: Learning from Visually Grounded Post-Training |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2604.05117 |