STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Yueqian, Wang, Yuxuan, Chen, Kai, Zhao, Dongyan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910290075975680
author Wang, Yueqian
Wang, Yuxuan
Chen, Kai
Zhao, Dongyan
author_facet Wang, Yueqian
Wang, Yuxuan
Chen, Kai
Zhao, Dongyan
contents Recently we have witnessed the rapid development of video question answering models. However, most models can only handle simple videos in terms of temporal reasoning, and their performance tends to drop when answering temporal-reasoning questions on long and informative videos. To tackle this problem we propose STAIR, a Spatial-Temporal Reasoning model with Auditable Intermediate Results for video question answering. STAIR is a neural module network, which contains a program generator to decompose a given question into a hierarchical combination of several sub-tasks, and a set of lightweight neural modules to complete each of these sub-tasks. Though neural module networks are already widely studied on image-text tasks, applying them to videos is a non-trivial task, as reasoning on videos requires different abilities. In this paper, we define a set of basic video-text sub-tasks for video question answering and design a set of lightweight modules to complete them. Different from most prior works, modules of STAIR return intermediate outputs specific to their intentions instead of always returning attention maps, which makes it easier to interpret and collaborate with pre-trained models. We also introduce intermediate supervision to make these intermediate outputs more accurate. We conduct extensive experiments on several video question answering datasets under various settings to show STAIR's performance, explainability, compatibility with pre-trained models, and applicability when program annotations are not available. Code: https://github.com/yellow-binary-tree/STAIR
format Preprint
id arxiv_https___arxiv_org_abs_2401_03901
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
Wang, Yueqian
Wang, Yuxuan
Chen, Kai
Zhao, Dongyan
Computer Vision and Pattern Recognition
Computation and Language
Recently we have witnessed the rapid development of video question answering models. However, most models can only handle simple videos in terms of temporal reasoning, and their performance tends to drop when answering temporal-reasoning questions on long and informative videos. To tackle this problem we propose STAIR, a Spatial-Temporal Reasoning model with Auditable Intermediate Results for video question answering. STAIR is a neural module network, which contains a program generator to decompose a given question into a hierarchical combination of several sub-tasks, and a set of lightweight neural modules to complete each of these sub-tasks. Though neural module networks are already widely studied on image-text tasks, applying them to videos is a non-trivial task, as reasoning on videos requires different abilities. In this paper, we define a set of basic video-text sub-tasks for video question answering and design a set of lightweight modules to complete them. Different from most prior works, modules of STAIR return intermediate outputs specific to their intentions instead of always returning attention maps, which makes it easier to interpret and collaborate with pre-trained models. We also introduce intermediate supervision to make these intermediate outputs more accurate. We conduct extensive experiments on several video question answering datasets under various settings to show STAIR's performance, explainability, compatibility with pre-trained models, and applicability when program annotations are not available. Code: https://github.com/yellow-binary-tree/STAIR
title STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2401.03901