ReasVQA: Advancing VideoQA with Imperfect Reasoning Process

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Jianxin, Meng, Xiaojun, Zhang, Huishuai, Wang, Yueqian, Wei, Jiansheng, Zhao, Dongyan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909464371658752
author Liang, Jianxin
Meng, Xiaojun
Zhang, Huishuai
Wang, Yueqian
Wei, Jiansheng
Zhao, Dongyan
author_facet Liang, Jianxin
Meng, Xiaojun
Zhang, Huishuai
Wang, Yueqian
Wei, Jiansheng
Zhao, Dongyan
contents Video Question Answering (VideoQA) is a challenging task that requires understanding complex visual and temporal relationships within videos to answer questions accurately. In this work, we introduce \textbf{ReasVQA} (Reasoning-enhanced Video Question Answering), a novel approach that leverages reasoning processes generated by Multimodal Large Language Models (MLLMs) to improve the performance of VideoQA models. Our approach consists of three phases: reasoning generation, reasoning refinement, and learning from reasoning. First, we generate detailed reasoning processes using additional MLLMs, and second refine them via a filtering step to ensure data quality. Finally, we use the reasoning data, which might be in an imperfect form, to guide the VideoQA model via multi-task learning, on how to interpret and answer questions based on a given video. We evaluate ReasVQA on three popular benchmarks, and our results establish new state-of-the-art performance with significant improvements of +2.9 on NExT-QA, +7.3 on STAR, and +5.9 on IntentQA. Our findings demonstrate the supervising benefits of integrating reasoning processes into VideoQA. Further studies validate each component of our method, also with different backbones and MLLMs, and again highlight the advantages of this simple but effective method. We offer a new perspective on enhancing VideoQA performance by utilizing advanced reasoning techniques, setting a new benchmark in this research field.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13536
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
Liang, Jianxin
Meng, Xiaojun
Zhang, Huishuai
Wang, Yueqian
Wei, Jiansheng
Zhao, Dongyan
Computer Vision and Pattern Recognition
Computation and Language
Video Question Answering (VideoQA) is a challenging task that requires understanding complex visual and temporal relationships within videos to answer questions accurately. In this work, we introduce \textbf{ReasVQA} (Reasoning-enhanced Video Question Answering), a novel approach that leverages reasoning processes generated by Multimodal Large Language Models (MLLMs) to improve the performance of VideoQA models. Our approach consists of three phases: reasoning generation, reasoning refinement, and learning from reasoning. First, we generate detailed reasoning processes using additional MLLMs, and second refine them via a filtering step to ensure data quality. Finally, we use the reasoning data, which might be in an imperfect form, to guide the VideoQA model via multi-task learning, on how to interpret and answer questions based on a given video. We evaluate ReasVQA on three popular benchmarks, and our results establish new state-of-the-art performance with significant improvements of +2.9 on NExT-QA, +7.3 on STAR, and +5.9 on IntentQA. Our findings demonstrate the supervising benefits of integrating reasoning processes into VideoQA. Further studies validate each component of our method, also with different backbones and MLLMs, and again highlight the advantages of this simple but effective method. We offer a new perspective on enhancing VideoQA performance by utilizing advanced reasoning techniques, setting a new benchmark in this research field.
title ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2501.13536