WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuanhan, Zhang, Kaichen, Li, Bo, Pu, Fanyi, Setiadharma, Christopher Arif, Yang, Jingkang, Liu, Ziwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929336305582080
author Zhang, Yuanhan
Zhang, Kaichen
Li, Bo
Pu, Fanyi
Setiadharma, Christopher Arif
Yang, Jingkang
Liu, Ziwei
author_facet Zhang, Yuanhan
Zhang, Kaichen
Li, Bo
Pu, Fanyi
Setiadharma, Christopher Arif
Yang, Jingkang
Liu, Ziwei
contents Multimodal information, together with our knowledge, help us to understand the complex and dynamic world. Large language models (LLM) and large multimodal models (LMM), however, still struggle to emulate this capability. In this paper, we present WorldQA, a video understanding dataset designed to push the boundaries of multimodal world models with three appealing properties: (1) Multimodal Inputs: The dataset comprises 1007 question-answer pairs and 303 videos, necessitating the analysis of both auditory and visual data for successful interpretation. (2) World Knowledge: We identify five essential types of world knowledge for question formulation. This approach challenges models to extend their capabilities beyond mere perception. (3) Long-Chain Reasoning: Our dataset introduces an average reasoning step of 4.45, notably surpassing other videoQA datasets. Furthermore, we introduce WorldRetriever, an agent designed to synthesize expert knowledge into a coherent reasoning chain, thereby facilitating accurate responses to WorldQA queries. Extensive evaluations of 13 prominent LLMs and LMMs reveal that WorldRetriever, although being the most effective model, achieved only 70% of humanlevel performance in multiple-choice questions. This finding highlights the necessity for further advancement in the reasoning and comprehension abilities of models. Our experiments also yield several key insights. For instance, while humans tend to perform better with increased frames, current LMMs, including WorldRetriever, show diminished performance under similar conditions. We hope that WorldQA,our methodology, and these insights could contribute to the future development of multimodal world models.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03272
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
Zhang, Yuanhan
Zhang, Kaichen
Li, Bo
Pu, Fanyi
Setiadharma, Christopher Arif
Yang, Jingkang
Liu, Ziwei
Computer Vision and Pattern Recognition
Multimodal information, together with our knowledge, help us to understand the complex and dynamic world. Large language models (LLM) and large multimodal models (LMM), however, still struggle to emulate this capability. In this paper, we present WorldQA, a video understanding dataset designed to push the boundaries of multimodal world models with three appealing properties: (1) Multimodal Inputs: The dataset comprises 1007 question-answer pairs and 303 videos, necessitating the analysis of both auditory and visual data for successful interpretation. (2) World Knowledge: We identify five essential types of world knowledge for question formulation. This approach challenges models to extend their capabilities beyond mere perception. (3) Long-Chain Reasoning: Our dataset introduces an average reasoning step of 4.45, notably surpassing other videoQA datasets. Furthermore, we introduce WorldRetriever, an agent designed to synthesize expert knowledge into a coherent reasoning chain, thereby facilitating accurate responses to WorldQA queries. Extensive evaluations of 13 prominent LLMs and LMMs reveal that WorldRetriever, although being the most effective model, achieved only 70% of humanlevel performance in multiple-choice questions. This finding highlights the necessity for further advancement in the reasoning and comprehension abilities of models. Our experiments also yield several key insights. For instance, while humans tend to perform better with increased frames, current LMMs, including WorldRetriever, show diminished performance under similar conditions. We hope that WorldQA,our methodology, and these insights could contribute to the future development of multimodal world models.
title WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.03272