CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Xingcheng, Guo, Hao, Song, Rui, Zimmer, Walter, Liu, Mingyu, Schamschurko, André, Cao, Hu, Knoll, Alois |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events
by: Zhou, Xingcheng, et al.
Published: (2024)
by: Zhou, Xingcheng, et al.
Published: (2024)
PointCompress3D: A Point Cloud Compression Framework for Roadside LiDARs in Intelligent Transportation Systems
by: Zimmer, Walter, et al.
Published: (2024)
by: Zimmer, Walter, et al.
Published: (2024)
SegRGB-X: General RGB-X Semantic Segmentation Model
by: Liu, Jiong, et al.
Published: (2026)
by: Liu, Jiong, et al.
Published: (2026)
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
by: Song, Zijie, et al.
Published: (2025)
by: Song, Zijie, et al.
Published: (2025)
RECSIP: REpeated Clustering of Scores Improving the Precision
by: Schamschurko, André, et al.
Published: (2025)
by: Schamschurko, André, et al.
Published: (2025)
VideoQA in the Era of LLMs: An Empirical Study
by: Xiao, Junbin, et al.
Published: (2024)
by: Xiao, Junbin, et al.
Published: (2024)
Vision Language Models in Autonomous Driving: A Survey and Outlook
by: Zhou, Xingcheng, et al.
Published: (2023)
by: Zhou, Xingcheng, et al.
Published: (2023)
TUMTraf EMOT: Event-Based Multi-Object Tracking Dataset and Baseline for Traffic Scenarios
by: Li, Mengyu, et al.
Published: (2025)
by: Li, Mengyu, et al.
Published: (2025)
GenAI-Driven Approach to RISC-V Supply Chain Exploration
by: Petrovic, Nenad, et al.
Published: (2026)
by: Petrovic, Nenad, et al.
Published: (2026)
TUMTraf V2X Cooperative Perception Dataset
by: Zimmer, Walter, et al.
Published: (2024)
by: Zimmer, Walter, et al.
Published: (2024)
Safety-Critical Learning for Long-Tail Events: The TUM Traffic Accident Dataset
by: Zimmer, Walter, et al.
Published: (2025)
by: Zimmer, Walter, et al.
Published: (2025)
WARM-3D: A Weakly-Supervised Sim2Real Domain Adaptation Framework for Roadside Monocular 3D Object Detection
by: Zhou, Xingcheng, et al.
Published: (2024)
by: Zhou, Xingcheng, et al.
Published: (2024)
Towards Vision Zero: The TUM Traffic Accid3nD Dataset
by: Zimmer, Walter, et al.
Published: (2025)
by: Zimmer, Walter, et al.
Published: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
VideoQA-SC: Adaptive Semantic Communication for Video Question Answering
by: Guo, Jiangyuan, et al.
Published: (2024)
by: Guo, Jiangyuan, et al.
Published: (2024)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
ENTER: Event Based Interpretable Reasoning for VideoQA
by: Ayyubi, Hammad, et al.
Published: (2025)
by: Ayyubi, Hammad, et al.
Published: (2025)
Reading Between the Lanes: Text VideoQA on the Road
by: Tom, George, et al.
Published: (2023)
by: Tom, George, et al.
Published: (2023)
A Survey on Autonomous Driving Datasets: Statistics, Annotation Quality, and a Future Outlook
by: Liu, Mingyu, et al.
Published: (2024)
by: Liu, Mingyu, et al.
Published: (2024)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
by: Hu, Yuhang, et al.
Published: (2025)
by: Hu, Yuhang, et al.
Published: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
Understanding Complexity in VideoQA via Visual Program Generation
by: Eyzaguirre, Cristobal, et al.
Published: (2025)
by: Eyzaguirre, Cristobal, et al.
Published: (2025)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
GraphRelate3D: Context-Dependent 3D Object Detection with Inter-Object Relationship Graphs
by: Liu, Mingyu, et al.
Published: (2024)
by: Liu, Mingyu, et al.
Published: (2024)
Are requirements really all you need? A case study of LLM-driven configuration code generation for automotive simulations
by: Lebioda, Krzysztof, et al.
Published: (2025)
by: Lebioda, Krzysztof, et al.
Published: (2025)
GenAI for Automotive Software Development: From Requirements to Wheels
by: Petrovic, Nenad, et al.
Published: (2025)
by: Petrovic, Nenad, et al.
Published: (2025)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems
by: He, Zhixian, et al.
Published: (2024)
by: He, Zhixian, et al.
Published: (2024)
Enhancing Highway Safety: Accident Detection on the A9 Test Stretch Using Roadside Sensors
by: Zimmer, Walter, et al.
Published: (2025)
by: Zimmer, Walter, et al.
Published: (2025)
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
by: Parikh, Chirag, et al.
Published: (2025)
by: Parikh, Chirag, et al.
Published: (2025)
CoDa-4DGS: Dynamic Gaussian Splatting with Context and Deformation Awareness for Autonomous Driving
by: Song, Rui, et al.
Published: (2025)
by: Song, Rui, et al.
Published: (2025)
Collaborative Semantic Occupancy Prediction with Hybrid Feature Fusion in Connected Automated Vehicles
by: Song, Rui, et al.
Published: (2024)
by: Song, Rui, et al.
Published: (2024)
Energy-Aware Imitation Learning for Steering Prediction Using Events and Frames
by: Cao, Hu, et al.
Published: (2026)
by: Cao, Hu, et al.
Published: (2026)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
by: Liang, Lili, et al.
Published: (2024)
by: Liang, Lili, et al.
Published: (2024)
BiSeg-SAM: Weakly-Supervised Post-Processing Framework for Boosting Binary Segmentation in Segment Anything Models
by: Su, Encheng, et al.
Published: (2025)
by: Su, Encheng, et al.
Published: (2025)
URNet: Uncertainty-aware Refinement Network for Event-based Stereo Depth Estimation
by: Cheng, Yifeng, et al.
Published: (2025)
by: Cheng, Yifeng, et al.
Published: (2025)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
Similar Items
-
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025) -
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026) -
GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events
by: Zhou, Xingcheng, et al.
Published: (2024) -
PointCompress3D: A Point Cloud Compression Framework for Roadside LiDARs in Intelligent Transportation Systems
by: Zimmer, Walter, et al.
Published: (2024) -
SegRGB-X: General RGB-X Semantic Segmentation Model
by: Liu, Jiong, et al.
Published: (2026)