TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Xingcheng, Larintzakis, Konstantinos, Guo, Hao, Zimmer, Walter, Liu, Mingyu, Cao, Hu, Zhang, Jiajie, Lakshminarasimhan, Venkatnarayanan, Strand, Leah, Knoll, Alois C.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913678042857472
author Zhou, Xingcheng
Larintzakis, Konstantinos
Guo, Hao
Zimmer, Walter
Liu, Mingyu
Cao, Hu
Zhang, Jiajie
Lakshminarasimhan, Venkatnarayanan
Strand, Leah
Knoll, Alois C.
author_facet Zhou, Xingcheng
Larintzakis, Konstantinos
Guo, Hao
Zimmer, Walter
Liu, Mingyu
Cao, Hu
Zhang, Jiajie
Lakshminarasimhan, Venkatnarayanan
Strand, Leah
Knoll, Alois C.
contents We present TUMTraffic-VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object captioning, and 5,700 object grounding annotations, encompassing diverse real-world conditions such as adverse weather and traffic anomalies. By incorporating tuple-based spatio-temporal object expressions, TUMTraffic-VideoQA unifies three essential tasks-multiple-choice video question answering, referred object captioning, and spatio-temporal object grounding-within a cohesive evaluation framework. We further introduce the TUMTraffic-Qwen baseline model, enhanced with visual token sampling strategies, providing valuable insights into the challenges of fine-grained spatio-temporal reasoning. Extensive experiments demonstrate the dataset's complexity, highlight the limitations of existing models, and position TUMTraffic-VideoQA as a robust foundation for advancing research in intelligent transportation systems. The dataset and benchmark are publicly available to facilitate further exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02449
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
Zhou, Xingcheng
Larintzakis, Konstantinos
Guo, Hao
Zimmer, Walter
Liu, Mingyu
Cao, Hu
Zhang, Jiajie
Lakshminarasimhan, Venkatnarayanan
Strand, Leah
Knoll, Alois C.
Computer Vision and Pattern Recognition
We present TUMTraffic-VideoQA, a novel dataset and benchmark designed for spatio-temporal video understanding in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object captioning, and 5,700 object grounding annotations, encompassing diverse real-world conditions such as adverse weather and traffic anomalies. By incorporating tuple-based spatio-temporal object expressions, TUMTraffic-VideoQA unifies three essential tasks-multiple-choice video question answering, referred object captioning, and spatio-temporal object grounding-within a cohesive evaluation framework. We further introduce the TUMTraffic-Qwen baseline model, enhanced with visual token sampling strategies, providing valuable insights into the challenges of fine-grained spatio-temporal reasoning. Extensive experiments demonstrate the dataset's complexity, highlight the limitations of existing models, and position TUMTraffic-VideoQA as a robust foundation for advancing research in intelligent transportation systems. The dataset and benchmark are publicly available to facilitate further exploration.
title TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.02449