Reading Between the Lanes: Text VideoQA on the Road

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tom, George, Mathew, Minesh, Garcia, Sergi, Karatzas, Dimosthenis, Jawahar, C. V.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909648700833792
author Tom, George
Mathew, Minesh
Garcia, Sergi
Karatzas, Dimosthenis
Jawahar, C. V.
author_facet Tom, George
Mathew, Minesh
Garcia, Sergi
Karatzas, Dimosthenis
Jawahar, C. V.
contents Text and signs around roads provide crucial information for drivers, vital for safe navigation and situational awareness. Scene text recognition in motion is a challenging problem, while textual cues typically appear for a short time span, and early detection at a distance is necessary. Systems that exploit such information to assist the driver should not only extract and incorporate visual and textual cues from the video stream but also reason over time. To address this issue, we introduce RoadTextVQA, a new dataset for the task of video question answering (VideoQA) in the context of driver assistance. RoadTextVQA consists of $3,222$ driving videos collected from multiple countries, annotated with $10,500$ questions, all based on text or road signs present in the driving videos. We assess the performance of state-of-the-art video question answering models on our RoadTextVQA dataset, highlighting the significant potential for improvement in this domain and the usefulness of the dataset in advancing research on in-vehicle support systems and text-aware multimodal question answering. The dataset is available at http://cvit.iiit.ac.in/research/projects/cvit-projects/roadtextvqa
format Preprint
id arxiv_https___arxiv_org_abs_2307_03948
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Reading Between the Lanes: Text VideoQA on the Road
Tom, George
Mathew, Minesh
Garcia, Sergi
Karatzas, Dimosthenis
Jawahar, C. V.
Computer Vision and Pattern Recognition
Text and signs around roads provide crucial information for drivers, vital for safe navigation and situational awareness. Scene text recognition in motion is a challenging problem, while textual cues typically appear for a short time span, and early detection at a distance is necessary. Systems that exploit such information to assist the driver should not only extract and incorporate visual and textual cues from the video stream but also reason over time. To address this issue, we introduce RoadTextVQA, a new dataset for the task of video question answering (VideoQA) in the context of driver assistance. RoadTextVQA consists of $3,222$ driving videos collected from multiple countries, annotated with $10,500$ questions, all based on text or road signs present in the driving videos. We assess the performance of state-of-the-art video question answering models on our RoadTextVQA dataset, highlighting the significant potential for improvement in this domain and the usefulness of the dataset in advancing research on in-vehicle support systems and text-aware multimodal question answering. The dataset is available at http://cvit.iiit.ac.in/research/projects/cvit-projects/roadtextvqa
title Reading Between the Lanes: Text VideoQA on the Road
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2307.03948