GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Xingcheng, Knoll, Alois C. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TUMTraf EMOT: Event-Based Multi-Object Tracking Dataset and Baseline for Traffic Scenarios
by: Li, Mengyu, et al.
Published: (2025)
by: Li, Mengyu, et al.
Published: (2025)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
Safety-Critical Learning for Long-Tail Events: The TUM Traffic Accident Dataset
by: Zimmer, Walter, et al.
Published: (2025)
by: Zimmer, Walter, et al.
Published: (2025)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
Towards Vision Zero: The TUM Traffic Accid3nD Dataset
by: Zimmer, Walter, et al.
Published: (2025)
by: Zimmer, Walter, et al.
Published: (2025)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
Vision Language Models in Autonomous Driving: A Survey and Outlook
by: Zhou, Xingcheng, et al.
Published: (2023)
by: Zhou, Xingcheng, et al.
Published: (2023)
TUMTraf V2X Cooperative Perception Dataset
by: Zimmer, Walter, et al.
Published: (2024)
by: Zimmer, Walter, et al.
Published: (2024)
SegRGB-X: General RGB-X Semantic Segmentation Model
by: Liu, Jiong, et al.
Published: (2026)
by: Liu, Jiong, et al.
Published: (2026)
URNet: Uncertainty-aware Refinement Network for Event-based Stereo Depth Estimation
by: Cheng, Yifeng, et al.
Published: (2025)
by: Cheng, Yifeng, et al.
Published: (2025)
PointCompress3D: A Point Cloud Compression Framework for Roadside LiDARs in Intelligent Transportation Systems
by: Zimmer, Walter, et al.
Published: (2024)
by: Zimmer, Walter, et al.
Published: (2024)
WARM-3D: A Weakly-Supervised Sim2Real Domain Adaptation Framework for Roadside Monocular 3D Object Detection
by: Zhou, Xingcheng, et al.
Published: (2024)
by: Zhou, Xingcheng, et al.
Published: (2024)
DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving
by: Kirchner, Sven, et al.
Published: (2025)
by: Kirchner, Sven, et al.
Published: (2025)
Hallucination Elimination and Semantic Enhancement Framework for Vision-Language Models in Traffic Scenarios
by: Fan, Jiaqi, et al.
Published: (2024)
by: Fan, Jiaqi, et al.
Published: (2024)
SEVD: Synthetic Event-based Vision Dataset for Ego and Fixed Traffic Perception
by: Aliminati, Manideep Reddy, et al.
Published: (2024)
by: Aliminati, Manideep Reddy, et al.
Published: (2024)
LFA: Layer Feature Attention for Run-Time Introspection of 2D Object Detectors in Automated Driving
by: Keser, Mert, et al.
Published: (2026)
by: Keser, Mert, et al.
Published: (2026)
Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events
by: Tami, Mohammad Abu, et al.
Published: (2024)
by: Tami, Mohammad Abu, et al.
Published: (2024)
A Survey on Autonomous Driving Datasets: Statistics, Annotation Quality, and a Future Outlook
by: Liu, Mingyu, et al.
Published: (2024)
by: Liu, Mingyu, et al.
Published: (2024)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
by: Ma, Jingtian, et al.
Published: (2025)
by: Ma, Jingtian, et al.
Published: (2025)
Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
by: Theodoridis, Nikos, et al.
Published: (2025)
by: Theodoridis, Nikos, et al.
Published: (2025)
eTraM: Event-based Traffic Monitoring Dataset
by: Verma, Aayush Atul, et al.
Published: (2024)
by: Verma, Aayush Atul, et al.
Published: (2024)
EMF: Event Meta Formers for Event-based Real-time Traffic Object Detection
by: Khan, Muhammad Ahmed Ullah, et al.
Published: (2025)
by: Khan, Muhammad Ahmed Ullah, et al.
Published: (2025)
Vision Technologies with Applications in Traffic Surveillance Systems: A Holistic Survey
by: Zhou, Wei, et al.
Published: (2024)
by: Zhou, Wei, et al.
Published: (2024)
Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
by: Masri, Sari, et al.
Published: (2025)
by: Masri, Sari, et al.
Published: (2025)
CATS-V2V: A Real-World Vehicle-to-Vehicle Cooperative Perception Dataset with Complex Adverse Traffic Scenarios
by: Li, Hangyu, et al.
Published: (2025)
by: Li, Hangyu, et al.
Published: (2025)
Benchmarking Vision Foundation Models for Input Monitoring in Autonomous Driving
by: Keser, Mert, et al.
Published: (2025)
by: Keser, Mert, et al.
Published: (2025)
The Urban Vision Hackathon Dataset and Models: Towards Image Annotations and Accurate Vision Models for Indian Traffic
by: Sharma, Akash, et al.
Published: (2025)
by: Sharma, Akash, et al.
Published: (2025)
Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis
by: Lin, Yu-Hsuan
Published: (2025)
by: Lin, Yu-Hsuan
Published: (2025)
Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
by: Lohner, Aaron, et al.
Published: (2024)
by: Lohner, Aaron, et al.
Published: (2024)
Enhancing Traffic Object Detection in Variable Illumination with RGB-Event Fusion
by: Liu, Zhanwen, et al.
Published: (2023)
by: Liu, Zhanwen, et al.
Published: (2023)
How Could Generative AI Support Compliance with the EU AI Act? A Review for Safe Automated Driving Perception
by: Keser, Mert, et al.
Published: (2024)
by: Keser, Mert, et al.
Published: (2024)
GraphRelate3D: Context-Dependent 3D Object Detection with Inter-Object Relationship Graphs
by: Liu, Mingyu, et al.
Published: (2024)
by: Liu, Mingyu, et al.
Published: (2024)
A Simple Framework Towards Vision-based Traffic Signal Control with Microscopic Simulation
by: He, Pan, et al.
Published: (2024)
by: He, Pan, et al.
Published: (2024)
TUMTraf Event: Calibration and Fusion Resulting in a Dataset for Roadside Event-Based and RGB Cameras
by: Creß, Christian, et al.
Published: (2024)
by: Creß, Christian, et al.
Published: (2024)
Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations
by: Garg, Sparsh, et al.
Published: (2025)
by: Garg, Sparsh, et al.
Published: (2025)
AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward
by: Han, Haonan, et al.
Published: (2024)
by: Han, Haonan, et al.
Published: (2024)
Revolutionizing Traffic Sign Recognition: Unveiling the Potential of Vision Transformers
by: Mingwin, Susano, et al.
Published: (2024)
by: Mingwin, Susano, et al.
Published: (2024)
EventGPT: Event Stream Understanding with Multimodal Large Language Models
by: Liu, Shaoyu, et al.
Published: (2024)
by: Liu, Shaoyu, et al.
Published: (2024)
Enhancing Highway Safety: Accident Detection on the A9 Test Stretch Using Roadside Sensors
by: Zimmer, Walter, et al.
Published: (2025)
by: Zimmer, Walter, et al.
Published: (2025)
Similar Items
-
TUMTraf EMOT: Event-Based Multi-Object Tracking Dataset and Baseline for Traffic Scenarios
by: Li, Mengyu, et al.
Published: (2025) -
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026) -
Safety-Critical Learning for Long-Tail Events: The TUM Traffic Accident Dataset
by: Zimmer, Walter, et al.
Published: (2025) -
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
by: Zhou, Xingcheng, et al.
Published: (2026) -
Towards Vision Zero: The TUM Traffic Accid3nD Dataset
by: Zimmer, Walter, et al.
Published: (2025)