VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Swetha, Sirnam, Gupta, Rohit, Kulkarni, Parth Parag, Shatwell, David G, Santiago, Jeffrey A Chan, Siddiqui, Nyle, Fioresi, Joseph, Shah, Mubarak
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910082840657920
author Swetha, Sirnam
Gupta, Rohit
Kulkarni, Parth Parag
Shatwell, David G
Santiago, Jeffrey A Chan
Siddiqui, Nyle
Fioresi, Joseph
Shah, Mubarak
author_facet Swetha, Sirnam
Gupta, Rohit
Kulkarni, Parth Parag
Shatwell, David G
Santiago, Jeffrey A Chan
Siddiqui, Nyle
Fioresi, Joseph
Shah, Mubarak
contents Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly focus on questions answerable through explicit visual content - actions, objects, and events - directly observable within individual frames or short clips. To truly understand videos as humans do, models must go beyond what is directly shown, inferring hidden relationships and contextual cues that are only implied across frames. Current benchmarks fail to capture this essential aspect of video understanding. To address this gap, we introduce VRR-QA, a benchmark for Visual Relational Reasoning Beyond Explicit Cues. We curate our benchmark from creative and cinematic videos such as movies, that deliberately employ storytelling techniques which omit direct depictions of certain events or relations, requiring viewers to infer them. VRR-QA comprises 1K meticulously expert-annotated QA pairs drawn from 1K creative video clips covering 15 genres across 7 decades of content, from both live-action and animated titles. Our extensive evaluations on 14 leading VideoQA models reveals consistent and significant performance degradation, underscoring their reliance on surface-level visual cues and highlighting the difficulty of implicit reasoning. Even the best model substantially underperforms human baselines with only 64% accuracy. Performance variations across models further illustrate the complexity and diversity of the challenges presented by VRR-QA. By releasing both dataset and data collection framework, VRR-QA establishes a rigorous, diverse, and reproducible testbed for advancing VideoQA: https://swetha5.github.io/ImplicitQA/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21742
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
Swetha, Sirnam
Gupta, Rohit
Kulkarni, Parth Parag
Shatwell, David G
Santiago, Jeffrey A Chan
Siddiqui, Nyle
Fioresi, Joseph
Shah, Mubarak
Computer Vision and Pattern Recognition
Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly focus on questions answerable through explicit visual content - actions, objects, and events - directly observable within individual frames or short clips. To truly understand videos as humans do, models must go beyond what is directly shown, inferring hidden relationships and contextual cues that are only implied across frames. Current benchmarks fail to capture this essential aspect of video understanding. To address this gap, we introduce VRR-QA, a benchmark for Visual Relational Reasoning Beyond Explicit Cues. We curate our benchmark from creative and cinematic videos such as movies, that deliberately employ storytelling techniques which omit direct depictions of certain events or relations, requiring viewers to infer them. VRR-QA comprises 1K meticulously expert-annotated QA pairs drawn from 1K creative video clips covering 15 genres across 7 decades of content, from both live-action and animated titles. Our extensive evaluations on 14 leading VideoQA models reveals consistent and significant performance degradation, underscoring their reliance on surface-level visual cues and highlighting the difficulty of implicit reasoning. Even the best model substantially underperforms human baselines with only 64% accuracy. Performance variations across models further illustrate the complexity and diversity of the challenges presented by VRR-QA. By releasing both dataset and data collection framework, VRR-QA establishes a rigorous, diverse, and reproducible testbed for advancing VideoQA: https://swetha5.github.io/ImplicitQA/.
title VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21742