Towards Fine-Grained Video Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Wei, Luo, Alan, Durante, Zane, Dash, Debadutta, Milstein, Arnold, Schulman, Kevin, Adeli, Ehsan, Fei-Fei, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929751165239296
author Dai, Wei
Luo, Alan
Durante, Zane
Dash, Debadutta
Milstein, Arnold
Schulman, Kevin
Adeli, Ehsan
Fei-Fei, Li
author_facet Dai, Wei
Luo, Alan
Durante, Zane
Dash, Debadutta
Milstein, Arnold
Schulman, Kevin
Adeli, Ehsan
Fei-Fei, Li
contents In the rapidly evolving domain of video understanding, Video Question Answering (VideoQA) remains a focal point. However, existing datasets exhibit gaps in temporal and spatial granularity, which consequently limits the capabilities of existing VideoQA methods. This paper introduces the Multi-Object Multi-Actor Question Answering (MOMA-QA) dataset, which is designed to address these shortcomings by emphasizing temporal localization, spatial relationship reasoning, and entity-centric queries. With ground truth scene graphs and temporal interval annotations, MOMA-QA is ideal for developing models for fine-grained video understanding. Furthermore, we present a novel video-language model, SGVLM, which incorporates a scene graph predictor, an efficient frame retriever, and a pre-trained large language model for temporal localization and fine-grained relationship understanding. Evaluations on MOMA-QA and other public datasets demonstrate the superior performance of our model, setting new benchmarks for VideoQA.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06820
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Fine-Grained Video Question Answering
Dai, Wei
Luo, Alan
Durante, Zane
Dash, Debadutta
Milstein, Arnold
Schulman, Kevin
Adeli, Ehsan
Fei-Fei, Li
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
In the rapidly evolving domain of video understanding, Video Question Answering (VideoQA) remains a focal point. However, existing datasets exhibit gaps in temporal and spatial granularity, which consequently limits the capabilities of existing VideoQA methods. This paper introduces the Multi-Object Multi-Actor Question Answering (MOMA-QA) dataset, which is designed to address these shortcomings by emphasizing temporal localization, spatial relationship reasoning, and entity-centric queries. With ground truth scene graphs and temporal interval annotations, MOMA-QA is ideal for developing models for fine-grained video understanding. Furthermore, we present a novel video-language model, SGVLM, which incorporates a scene graph predictor, an efficient frame retriever, and a pre-trained large language model for temporal localization and fine-grained relationship understanding. Evaluations on MOMA-QA and other public datasets demonstrate the superior performance of our model, setting new benchmarks for VideoQA.
title Towards Fine-Grained Video Question Answering
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.06820