Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Vishal, Joseph Raj, Basina, Divesh, Choudhary, Aarya, Chakravarthi, Bharatesh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InterAct-Video: Reasoning-Rich Video QA for Urban Traffic
by: Vishal, Joseph Raj, et al.
Published: (2025)
by: Vishal, Joseph Raj, et al.
Published: (2025)
KAT to KANs: A Review of Kolmogorov-Arnold Networks and the Neural Leap Forward
by: Basina, Divesh, et al.
Published: (2024)
by: Basina, Divesh, et al.
Published: (2024)
MC-BEVRO: Multi-Camera Bird Eye View Road Occupancy Detection for Traffic Monitoring
by: Vaghela, Arpitsinh, et al.
Published: (2025)
by: Vaghela, Arpitsinh, et al.
Published: (2025)
UDVideoQA: A Traffic Video Question Answering Dataset for Multi-Object Spatio-Temporal Reasoning in Urban Dynamics
by: Vishal, Joseph Raj, et al.
Published: (2026)
by: Vishal, Joseph Raj, et al.
Published: (2026)
Scale-Aware Vision-Language Adaptation for Extreme Far-Distance Video Person Re-identification
by: Rajbhandari, Ashwat, et al.
Published: (2026)
by: Rajbhandari, Ashwat, et al.
Published: (2026)
SKoPe3D: A Synthetic Dataset for Vehicle Keypoint Perception in 3D from Traffic Monitoring Cameras
by: Pahadia, Himanshu, et al.
Published: (2023)
by: Pahadia, Himanshu, et al.
Published: (2023)
eTraM: Event-based Traffic Monitoring Dataset
by: Verma, Aayush Atul, et al.
Published: (2024)
by: Verma, Aayush Atul, et al.
Published: (2024)
eSkiTB: A Synthetic Event-based Dataset for Tracking Skiers
by: Vinod, Krishna, et al.
Published: (2026)
by: Vinod, Krishna, et al.
Published: (2026)
How Real is CARLAs Dynamic Vision Sensor? A Study on the Sim-to-Real Gap in Traffic Object Detection
by: Tan, Kaiyuan, et al.
Published: (2025)
by: Tan, Kaiyuan, et al.
Published: (2025)
SEPose: A Synthetic Event-based Human Pose Estimation Dataset for Pedestrian Monitoring
by: Chanda, Kaustav, et al.
Published: (2025)
by: Chanda, Kaustav, et al.
Published: (2025)
Event Quality Score (EQS): Assessing the Realism of Simulated Event Camera Streams via Distances in Latent Space
by: Chanda, Kaustav, et al.
Published: (2025)
by: Chanda, Kaustav, et al.
Published: (2025)
Event-based Graph Representation with Spatial and Motion Vectors for Asynchronous Object Detection
by: Verma, Aayush Atul, et al.
Published: (2025)
by: Verma, Aayush Atul, et al.
Published: (2025)
Recent Event Camera Innovations: A Survey
by: Chakravarthi, Bharatesh, et al.
Published: (2024)
by: Chakravarthi, Bharatesh, et al.
Published: (2024)
SEVD: Synthetic Event-based Vision Dataset for Ego and Fixed Traffic Perception
by: Aliminati, Manideep Reddy, et al.
Published: (2024)
by: Aliminati, Manideep Reddy, et al.
Published: (2024)
SEBVS: Synthetic Event-based Visual Servoing for Robot Navigation and Manipulation
by: Vinod, Krishna, et al.
Published: (2025)
by: Vinod, Krishna, et al.
Published: (2025)
HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models
by: Bai, Xiangyu, et al.
Published: (2026)
by: Bai, Xiangyu, et al.
Published: (2026)
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
AdaFuse-Det: Adaptive Cross-Modal Fusion of Event Cameras for Robust Object Detection in Low-Light RGB Imagery
by: Imandi, Raju, et al.
Published: (2026)
by: Imandi, Raju, et al.
Published: (2026)
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
by: Pal, Aniket, et al.
Published: (2024)
by: Pal, Aniket, et al.
Published: (2024)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
Roundabout Dilemma Zone Data Mining and Forecasting with Trajectory Prediction and Graph Neural Networks
by: Satish, Manthan Chelenahalli, et al.
Published: (2024)
by: Satish, Manthan Chelenahalli, et al.
Published: (2024)
Admitting Ignorance Helps the Video Question Answering Models to Answer
by: Li, Haopeng, et al.
Published: (2025)
by: Li, Haopeng, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Agentic Keyframe Search for Video Question Answering
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
Narrative Aligned Long Form Video Question Answering
by: Jain, Rahul, et al.
Published: (2026)
by: Jain, Rahul, et al.
Published: (2026)
ViLA: Efficient Video-Language Alignment for Video Question Answering
by: Wang, Xijun, et al.
Published: (2023)
by: Wang, Xijun, et al.
Published: (2023)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025)
by: Di, Shangzhe, et al.
Published: (2025)
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
by: Oh, Ju-Young
Published: (2025)
by: Oh, Ju-Young
Published: (2025)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
by: Song, Enxin, et al.
Published: (2024)
by: Song, Enxin, et al.
Published: (2024)
Camera Perspective Transformation to Bird's Eye View via Spatial Transformer Model for Road Intersection Monitoring
by: Prajapati, Rukesh, et al.
Published: (2024)
by: Prajapati, Rukesh, et al.
Published: (2024)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
by: Nie, Yuxiang, et al.
Published: (2025)
by: Nie, Yuxiang, et al.
Published: (2025)
RoadscapesQA: A Multitask, Multimodal Dataset for Visual Question Answering on Indian Roads
by: Iyer, Vijayasri, et al.
Published: (2026)
by: Iyer, Vijayasri, et al.
Published: (2026)
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Descriptor: Distance-Annotated Traffic Perception Question Answering (DTPQA)
by: Theodoridis, Nikos, et al.
Published: (2025)
by: Theodoridis, Nikos, et al.
Published: (2025)
UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks
by: Nguyen, Jason, et al.
Published: (2026)
by: Nguyen, Jason, et al.
Published: (2026)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
by: Fernando, Basura, et al.
Published: (2025)
by: Fernando, Basura, et al.
Published: (2025)
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
by: Kugo, Noriyuki, et al.
Published: (2024)
by: Kugo, Noriyuki, et al.
Published: (2024)
Similar Items
-
InterAct-Video: Reasoning-Rich Video QA for Urban Traffic
by: Vishal, Joseph Raj, et al.
Published: (2025) -
KAT to KANs: A Review of Kolmogorov-Arnold Networks and the Neural Leap Forward
by: Basina, Divesh, et al.
Published: (2024) -
MC-BEVRO: Multi-Camera Bird Eye View Road Occupancy Detection for Traffic Monitoring
by: Vaghela, Arpitsinh, et al.
Published: (2025) -
UDVideoQA: A Traffic Video Question Answering Dataset for Multi-Object Spatio-Temporal Reasoning in Urban Dynamics
by: Vishal, Joseph Raj, et al.
Published: (2026) -
Scale-Aware Vision-Language Adaptation for Extreme Far-Distance Video Person Re-identification
by: Rajbhandari, Ashwat, et al.
Published: (2026)