Vision Language Models in Autonomous Driving: A Survey and Outlook
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Xingcheng, Liu, Mingyu, Yurtsever, Ekim, Zagar, Bare Luka, Zimmer, Walter, Cao, Hu, Knoll, Alois C. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey on Autonomous Driving Datasets: Statistics, Annotation Quality, and a Future Outlook
by: Liu, Mingyu, et al.
Published: (2024)
by: Liu, Mingyu, et al.
Published: (2024)
GraphRelate3D: Context-Dependent 3D Object Detection with Inter-Object Relationship Graphs
by: Liu, Mingyu, et al.
Published: (2024)
by: Liu, Mingyu, et al.
Published: (2024)
GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events
by: Zhou, Xingcheng, et al.
Published: (2024)
by: Zhou, Xingcheng, et al.
Published: (2024)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
PointCompress3D: A Point Cloud Compression Framework for Roadside LiDARs in Intelligent Transportation Systems
by: Zimmer, Walter, et al.
Published: (2024)
by: Zimmer, Walter, et al.
Published: (2024)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
SegRGB-X: General RGB-X Semantic Segmentation Model
by: Liu, Jiong, et al.
Published: (2026)
by: Liu, Jiong, et al.
Published: (2026)
WARM-3D: A Weakly-Supervised Sim2Real Domain Adaptation Framework for Roadside Monocular 3D Object Detection
by: Zhou, Xingcheng, et al.
Published: (2024)
by: Zhou, Xingcheng, et al.
Published: (2024)
CoDa-4DGS: Dynamic Gaussian Splatting with Context and Deformation Awareness for Autonomous Driving
by: Song, Rui, et al.
Published: (2025)
by: Song, Rui, et al.
Published: (2025)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
TUMTraf V2X Cooperative Perception Dataset
by: Zimmer, Walter, et al.
Published: (2024)
by: Zimmer, Walter, et al.
Published: (2024)
TUMTraf EMOT: Event-Based Multi-Object Tracking Dataset and Baseline for Traffic Scenarios
by: Li, Mengyu, et al.
Published: (2025)
by: Li, Mengyu, et al.
Published: (2025)
DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving
by: Kirchner, Sven, et al.
Published: (2025)
by: Kirchner, Sven, et al.
Published: (2025)
Benchmarking Vision Foundation Models for Input Monitoring in Autonomous Driving
by: Keser, Mert, et al.
Published: (2025)
by: Keser, Mert, et al.
Published: (2025)
Towards Vision Zero: The TUM Traffic Accid3nD Dataset
by: Zimmer, Walter, et al.
Published: (2025)
by: Zimmer, Walter, et al.
Published: (2025)
LFA: Layer Feature Attention for Run-Time Introspection of 2D Object Detectors in Automated Driving
by: Keser, Mert, et al.
Published: (2026)
by: Keser, Mert, et al.
Published: (2026)
Enhancing Highway Safety: Accident Detection on the A9 Test Stretch Using Roadside Sensors
by: Zimmer, Walter, et al.
Published: (2025)
by: Zimmer, Walter, et al.
Published: (2025)
Safety-Critical Learning for Long-Tail Events: The TUM Traffic Accident Dataset
by: Zimmer, Walter, et al.
Published: (2025)
by: Zimmer, Walter, et al.
Published: (2025)
URNet: Uncertainty-aware Refinement Network for Event-based Stereo Depth Estimation
by: Cheng, Yifeng, et al.
Published: (2025)
by: Cheng, Yifeng, et al.
Published: (2025)
Collaborative Semantic Occupancy Prediction with Hybrid Feature Fusion in Connected Automated Vehicles
by: Song, Rui, et al.
Published: (2024)
by: Song, Rui, et al.
Published: (2024)
Energy-Aware Imitation Learning for Steering Prediction Using Events and Frames
by: Cao, Hu, et al.
Published: (2026)
by: Cao, Hu, et al.
Published: (2026)
How Could Generative AI Support Compliance with the EU AI Act? A Review for Safe Automated Driving Perception
by: Keser, Mert, et al.
Published: (2024)
by: Keser, Mert, et al.
Published: (2024)
BiSeg-SAM: Weakly-Supervised Post-Processing Framework for Boosting Binary Segmentation in Segment Anything Models
by: Su, Encheng, et al.
Published: (2025)
by: Su, Encheng, et al.
Published: (2025)
A Unified Perception-Language-Action Framework for Adaptive Autonomous Driving
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
by: Jiang, Zebin, et al.
Published: (2025)
by: Jiang, Zebin, et al.
Published: (2025)
Visual Adversarial Attack on Vision-Language Models for Autonomous Driving
by: Zhang, Tianyuan, et al.
Published: (2024)
by: Zhang, Tianyuan, et al.
Published: (2024)
A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
by: Liu, Daizong, et al.
Published: (2024)
by: Liu, Daizong, et al.
Published: (2024)
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
by: Tian, Xiaoyu, et al.
Published: (2024)
by: Tian, Xiaoyu, et al.
Published: (2024)
Spatial-aware Vision Language Model for Autonomous Driving
by: Wei, Weijie, et al.
Published: (2025)
by: Wei, Weijie, et al.
Published: (2025)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
by: Zhao, Zongchuang, et al.
Published: (2025)
by: Zhao, Zongchuang, et al.
Published: (2025)
USC: Uncompromising Spatial Constraints for Safety-Oriented 3D Object Detectors in Autonomous Driving
by: Liao, Brian Hsuan-Cheng, et al.
Published: (2022)
by: Liao, Brian Hsuan-Cheng, et al.
Published: (2022)
MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving
by: Sun, Bin, et al.
Published: (2025)
by: Sun, Bin, et al.
Published: (2025)
VLP: Vision Language Planning for Autonomous Driving
by: Pan, Chenbin, et al.
Published: (2024)
by: Pan, Chenbin, et al.
Published: (2024)
DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
by: Diao, Muxi, et al.
Published: (2025)
by: Diao, Muxi, et al.
Published: (2025)
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models
by: Zhou, Xirui, et al.
Published: (2025)
by: Zhou, Xirui, et al.
Published: (2025)
TUMTraf Event: Calibration and Fusion Resulting in a Dataset for Roadside Event-Based and RGB Cameras
by: Creß, Christian, et al.
Published: (2024)
by: Creß, Christian, et al.
Published: (2024)
Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook
by: Bao, Muyi, et al.
Published: (2025)
by: Bao, Muyi, et al.
Published: (2025)
Similar Items
-
A Survey on Autonomous Driving Datasets: Statistics, Annotation Quality, and a Future Outlook
by: Liu, Mingyu, et al.
Published: (2024) -
GraphRelate3D: Context-Dependent 3D Object Detection with Inter-Object Relationship Graphs
by: Liu, Mingyu, et al.
Published: (2024) -
GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events
by: Zhou, Xingcheng, et al.
Published: (2024) -
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026) -
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
by: Zhou, Xingcheng, et al.
Published: (2025)