Reading Between the Lanes: Text VideoQA on the Road
Fuente:
arXiv
Saved in:
| Main Authors: | Tom, George, Mathew, Minesh, Garcia, Sergi, Karatzas, Dimosthenis, Jawahar, C. V. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Deployable OCR models for Indic languages
by: Mathew, Minesh, et al.
Published: (2022)
by: Mathew, Minesh, et al.
Published: (2022)
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
by: Gomez, Lluis, et al.
Published: (2014)
by: Gomez, Lluis, et al.
Published: (2014)
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
by: Parikh, Chirag, et al.
Published: (2025)
by: Parikh, Chirag, et al.
Published: (2025)
Reading in the Dark: Low-light Scene Text Recognition
by: Fu, Xuanshuo, et al.
Published: (2026)
by: Fu, Xuanshuo, et al.
Published: (2026)
VideoQA in the Era of LLMs: An Empirical Study
by: Xiao, Junbin, et al.
Published: (2024)
by: Xiao, Junbin, et al.
Published: (2024)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
Understanding Complexity in VideoQA via Visual Program Generation
by: Eyzaguirre, Cristobal, et al.
Published: (2025)
by: Eyzaguirre, Cristobal, et al.
Published: (2025)
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
by: Mishra, Pritam, et al.
Published: (2025)
by: Mishra, Pritam, et al.
Published: (2025)
ENTER: Event Based Interpretable Reasoning for VideoQA
by: Ayyubi, Hammad, et al.
Published: (2025)
by: Ayyubi, Hammad, et al.
Published: (2025)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems
by: He, Zhixian, et al.
Published: (2024)
by: He, Zhixian, et al.
Published: (2024)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
by: Mishra, Pritam, et al.
Published: (2026)
by: Mishra, Pritam, et al.
Published: (2026)
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
by: Song, Zijie, et al.
Published: (2025)
by: Song, Zijie, et al.
Published: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
by: Liang, Lili, et al.
Published: (2024)
by: Liang, Lili, et al.
Published: (2024)
VideoQA-SC: Adaptive Semantic Communication for Video Question Answering
by: Guo, Jiangyuan, et al.
Published: (2024)
by: Guo, Jiangyuan, et al.
Published: (2024)
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
by: Hu, Yuhang, et al.
Published: (2025)
by: Hu, Yuhang, et al.
Published: (2025)
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
by: Pintore, Marco, et al.
Published: (2025)
by: Pintore, Marco, et al.
Published: (2025)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
by: M, Megha Mariam K., et al.
Published: (2026)
by: M, Megha Mariam K., et al.
Published: (2026)
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
by: Ortega, Marc Serra, et al.
Published: (2025)
by: Ortega, Marc Serra, et al.
Published: (2025)
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
by: Li, Zongmin, et al.
Published: (2026)
by: Li, Zongmin, et al.
Published: (2026)
IndicSTR12: A Dataset for Indic Scene Text Recognition
by: Lunia, Harsh, et al.
Published: (2024)
by: Lunia, Harsh, et al.
Published: (2024)
Source-free Video Domain Adaptation by Learning from Noisy Labels
by: Dasgupta, Avijit, et al.
Published: (2023)
by: Dasgupta, Avijit, et al.
Published: (2023)
Image-text matching for large-scale book collections
by: Llabrés, Artemis, et al.
Published: (2024)
by: Llabrés, Artemis, et al.
Published: (2024)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2025)
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
Similar Items
-
Towards Deployable OCR models for Indic languages
by: Mathew, Minesh, et al.
Published: (2022) -
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
by: Pal, Aniket, et al.
Published: (2025) -
A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
by: Gomez, Lluis, et al.
Published: (2014) -
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
by: Parikh, Chirag, et al.
Published: (2025) -
Reading in the Dark: Low-light Scene Text Recognition
by: Fu, Xuanshuo, et al.
Published: (2026)