An Empirical Study for Representations of Videos in Video Question Answering via MLLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Zhi, Wang, Yanan, Niu, Hao, Vizcarra, Julio, Taya, Masato
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912646020726784
author Li, Zhi
Wang, Yanan
Niu, Hao
Vizcarra, Julio
Taya, Masato
author_facet Li, Zhi
Wang, Yanan
Niu, Hao
Vizcarra, Julio
Taya, Masato
contents Multimodal large language models have recently achieved remarkable progress in video question answering (VideoQA) by jointly processing visual, textual, and audio information. However, it remains unclear which video representations are most effective for MLLMs, and how different modalities balance task accuracy against computational efficiency. In this work, we present a comprehensive empirical study of video representation methods for VideoQA with MLLMs. We systematically evaluate single modality inputs question only, subtitles, visual frames, and audio signals as well as multimodal combinations, on two widely used benchmarks: VideoMME and LongVideoBench. Our results show that visual frames substantially enhance accuracy but impose heavy costs in GPU memory and inference latency, while subtitles provide a lightweight yet effective alternative, particularly for long videos. These findings highlight clear trade-offs between effectiveness and efficiency and provide practical insights for designing resource-aware MLLM-based VideoQA systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12299
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
Li, Zhi
Wang, Yanan
Niu, Hao
Vizcarra, Julio
Taya, Masato
Information Retrieval
I.2.10
Multimodal large language models have recently achieved remarkable progress in video question answering (VideoQA) by jointly processing visual, textual, and audio information. However, it remains unclear which video representations are most effective for MLLMs, and how different modalities balance task accuracy against computational efficiency. In this work, we present a comprehensive empirical study of video representation methods for VideoQA with MLLMs. We systematically evaluate single modality inputs question only, subtitles, visual frames, and audio signals as well as multimodal combinations, on two widely used benchmarks: VideoMME and LongVideoBench. Our results show that visual frames substantially enhance accuracy but impose heavy costs in GPU memory and inference latency, while subtitles provide a lightweight yet effective alternative, particularly for long videos. These findings highlight clear trade-offs between effectiveness and efficiency and provide practical insights for designing resource-aware MLLM-based VideoQA systems.
title An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
topic Information Retrieval
I.2.10
url https://arxiv.org/abs/2510.12299