Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Amoroso, Roberto, Zhang, Gengyuan, Koner, Rajat, Baraldi, Lorenzo, Cucchiara, Rita, Tresp, Volker
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929648731947008
author Amoroso, Roberto
Zhang, Gengyuan
Koner, Rajat
Baraldi, Lorenzo
Cucchiara, Rita
Tresp, Volker
author_facet Amoroso, Roberto
Zhang, Gengyuan
Koner, Rajat
Baraldi, Lorenzo
Cucchiara, Rita
Tresp, Volker
contents Video Question Answering (Video QA) is a challenging video understanding task that requires models to comprehend entire videos, identify the most relevant information based on contextual cues from a given question, and reason accurately to provide answers. Recent advancements in Multimodal Large Language Models (MLLMs) have transformed video QA by leveraging their exceptional commonsense reasoning capabilities. This progress is largely driven by the effective alignment between visual data and the language space of MLLMs. However, for video QA, an additional space-time alignment poses a considerable challenge for extracting question-relevant information across frames. In this work, we investigate diverse temporal modeling techniques to integrate with MLLMs, aiming to achieve question-guided temporal modeling that leverages pre-trained visual and textual alignment in MLLMs. We propose T-Former, a novel temporal modeling method that creates a question-guided temporal bridge between frame-wise visual perception and the reasoning capabilities of LLMs. Our evaluation across multiple video QA benchmarks demonstrates that T-Former competes favorably with existing temporal modeling approaches and aligns with recent advancements in video QA.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19304
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
Amoroso, Roberto
Zhang, Gengyuan
Koner, Rajat
Baraldi, Lorenzo
Cucchiara, Rita
Tresp, Volker
Computer Vision and Pattern Recognition
Video Question Answering (Video QA) is a challenging video understanding task that requires models to comprehend entire videos, identify the most relevant information based on contextual cues from a given question, and reason accurately to provide answers. Recent advancements in Multimodal Large Language Models (MLLMs) have transformed video QA by leveraging their exceptional commonsense reasoning capabilities. This progress is largely driven by the effective alignment between visual data and the language space of MLLMs. However, for video QA, an additional space-time alignment poses a considerable challenge for extracting question-relevant information across frames. In this work, we investigate diverse temporal modeling techniques to integrate with MLLMs, aiming to achieve question-guided temporal modeling that leverages pre-trained visual and textual alignment in MLLMs. We propose T-Former, a novel temporal modeling method that creates a question-guided temporal bridge between frame-wise visual perception and the reasoning capabilities of LLMs. Our evaluation across multiple video QA benchmarks demonstrates that T-Former competes favorably with existing temporal modeling approaches and aligns with recent advancements in video QA.
title Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.19304