Bridging Vision Language Models and Symbolic Grounding for Video Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Haodi, Pathak, Vyom, Wang, Daisy Zhe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916950033039360
author Ma, Haodi
Pathak, Vyom
Wang, Daisy Zhe
author_facet Ma, Haodi
Pathak, Vyom
Wang, Daisy Zhe
contents Video Question Answering (VQA) requires models to reason over spatial, temporal, and causal cues in videos. Recent vision language models (VLMs) achieve strong results but often rely on shallow correlations, leading to weak temporal grounding and limited interpretability. We study symbolic scene graphs (SGs) as intermediate grounding signals for VQA. SGs provide structured object-relation representations that complement VLMs holistic reasoning. We introduce SG-VLM, a modular framework that integrates frozen VLMs with scene graph grounding via prompting and visual localization. Across three benchmarks (NExT-QA, iVQA, ActivityNet-QA) and multiple VLMs (QwenVL, InternVL), SG-VLM improves causal and temporal reasoning and outperforms prior baselines, though gains over strong VLMs are limited. These findings highlight both the promise and current limitations of symbolic grounding, and offer guidance for future hybrid VLM-symbolic approaches in video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
Ma, Haodi
Pathak, Vyom
Wang, Daisy Zhe
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Video Question Answering (VQA) requires models to reason over spatial, temporal, and causal cues in videos. Recent vision language models (VLMs) achieve strong results but often rely on shallow correlations, leading to weak temporal grounding and limited interpretability. We study symbolic scene graphs (SGs) as intermediate grounding signals for VQA. SGs provide structured object-relation representations that complement VLMs holistic reasoning. We introduce SG-VLM, a modular framework that integrates frozen VLMs with scene graph grounding via prompting and visual localization. Across three benchmarks (NExT-QA, iVQA, ActivityNet-QA) and multiple VLMs (QwenVL, InternVL), SG-VLM improves causal and temporal reasoning and outperforms prior baselines, though gains over strong VLMs are limited. These findings highlight both the promise and current limitations of symbolic grounding, and offer guidance for future hybrid VLM-symbolic approaches in video understanding.
title Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.11862