Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Tieyuan, Liu, Huabin, Wang, Yi, Gan, Chaofan, Lyu, Mingxi, Qin, Ziran, Li, Shijie, Shen, Liquan, Hou, Junhui, Wang, Zheng, Lin, Weiyao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912757085896704
author Chen, Tieyuan
Liu, Huabin
Wang, Yi
Gan, Chaofan
Lyu, Mingxi
Qin, Ziran
Li, Shijie
Shen, Liquan
Hou, Junhui
Wang, Zheng
Lin, Weiyao
author_facet Chen, Tieyuan
Liu, Huabin
Wang, Yi
Gan, Chaofan
Lyu, Mingxi
Qin, Ziran
Li, Shijie
Shen, Liquan
Hou, Junhui
Wang, Zheng
Lin, Weiyao
contents Video Question Answering (VideoQA) aims to answer natural language questions based on the given video, with prior work primarily focusing on identifying the duration of relevant segments, referred to as explicit visual evidence. However, explicit visual evidence is not always directly available, particularly when questions target symbolic meanings or deeper intentions, leading to significant performance degradation. To fill this gap, we introduce a novel task and dataset, $\textbf{I}$mplicit $\textbf{V}$ideo $\textbf{Q}$uestion $\textbf{A}$nswering (I-VQA), which focuses on answering questions in scenarios where explicit visual evidence is inaccessible. Given an implicit question and its corresponding video, I-VQA requires answering based on the contextual visual cues present within the video. To tackle I-VQA, we propose a novel reasoning framework, IRM (Implicit Reasoning Model), incorporating dual-stream modeling of contextual actions and intent clues as implicit reasoning chains. IRM comprises the Action-Intent Module (AIM) and the Visual Enhancement Module (VEM). AIM deduces and preserves question-related dual clues by generating clue candidates and performing relation deduction. VEM enhances contextual visual representation by leveraging key contextual clues. Extensive experiments validate the effectiveness of our IRM in I-VQA tasks, outperforming GPT-4o, OpenAI-o3, and fine-tuned VideoChat2 by $0.76\%$, $1.37\%$, and $4.87\%$, respectively. Additionally, IRM performs SOTA on similar implicit advertisement understanding and future prediction in traffic-VQA. Datasets and codes are available for double-blind review in anonymous repo: https://github.com/tychen-SJTU/Implicit-VideoQA.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07811
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
Chen, Tieyuan
Liu, Huabin
Wang, Yi
Gan, Chaofan
Lyu, Mingxi
Qin, Ziran
Li, Shijie
Shen, Liquan
Hou, Junhui
Wang, Zheng
Lin, Weiyao
Computer Vision and Pattern Recognition
Video Question Answering (VideoQA) aims to answer natural language questions based on the given video, with prior work primarily focusing on identifying the duration of relevant segments, referred to as explicit visual evidence. However, explicit visual evidence is not always directly available, particularly when questions target symbolic meanings or deeper intentions, leading to significant performance degradation. To fill this gap, we introduce a novel task and dataset, $\textbf{I}$mplicit $\textbf{V}$ideo $\textbf{Q}$uestion $\textbf{A}$nswering (I-VQA), which focuses on answering questions in scenarios where explicit visual evidence is inaccessible. Given an implicit question and its corresponding video, I-VQA requires answering based on the contextual visual cues present within the video. To tackle I-VQA, we propose a novel reasoning framework, IRM (Implicit Reasoning Model), incorporating dual-stream modeling of contextual actions and intent clues as implicit reasoning chains. IRM comprises the Action-Intent Module (AIM) and the Visual Enhancement Module (VEM). AIM deduces and preserves question-related dual clues by generating clue candidates and performing relation deduction. VEM enhances contextual visual representation by leveraging key contextual clues. Extensive experiments validate the effectiveness of our IRM in I-VQA tasks, outperforming GPT-4o, OpenAI-o3, and fine-tuned VideoChat2 by $0.76\%$, $1.37\%$, and $4.87\%$, respectively. Additionally, IRM performs SOTA on similar implicit advertisement understanding and future prediction in traffic-VQA. Datasets and codes are available for double-blind review in anonymous repo: https://github.com/tychen-SJTU/Implicit-VideoQA.
title Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.07811