Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Aodi, Luo, Xubo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
SpaceSense-Bench: A Large-Scale Multi-Modal Benchmark for Spacecraft Perception and Pose Estimation
von: Wu, Aodi, et al.
Veröffentlicht: (2026)
von: Wu, Aodi, et al.
Veröffentlicht: (2026)
Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
von: Yu, Seungjun, et al.
Veröffentlicht: (2025)
von: Yu, Seungjun, et al.
Veröffentlicht: (2025)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
von: Mei, Xiaodong, et al.
Veröffentlicht: (2026)
von: Mei, Xiaodong, et al.
Veröffentlicht: (2026)
INSIGHT: Enhancing Autonomous Driving Safety through Vision-Language Models on Context-Aware Hazard Detection and Edge Case Evaluation
von: Chen, Dianwei, et al.
Veröffentlicht: (2025)
von: Chen, Dianwei, et al.
Veröffentlicht: (2025)
Learning Vision-Language-Action World Models for Autonomous Driving
von: Wang, Guoqing, et al.
Veröffentlicht: (2026)
von: Wang, Guoqing, et al.
Veröffentlicht: (2026)
Vision Language Models in Autonomous Driving: A Survey and Outlook
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2023)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2023)
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
von: Ogezi, Michael, et al.
Veröffentlicht: (2025)
von: Ogezi, Michael, et al.
Veröffentlicht: (2025)
HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios
von: Wang, Daming, et al.
Veröffentlicht: (2025)
von: Wang, Daming, et al.
Veröffentlicht: (2025)
A Survey on Vision-Language-Action Models for Autonomous Driving
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
Prompting Large Vision-Language Models for Compositional Reasoning
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
von: Wu, Junfei, et al.
Veröffentlicht: (2025)
von: Wu, Junfei, et al.
Veröffentlicht: (2025)
AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
von: Huang, Lianming, et al.
Veröffentlicht: (2025)
von: Huang, Lianming, et al.
Veröffentlicht: (2025)
Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving
von: Zhang, Enming, et al.
Veröffentlicht: (2025)
von: Zhang, Enming, et al.
Veröffentlicht: (2025)
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
von: Wang, Lu, et al.
Veröffentlicht: (2025)
von: Wang, Lu, et al.
Veröffentlicht: (2025)
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
Hybrid Reasoning Based on Large Language Models for Autonomous Car Driving
von: Azarafza, Mehdi, et al.
Veröffentlicht: (2024)
von: Azarafza, Mehdi, et al.
Veröffentlicht: (2024)
Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
von: Lian, Shijie, et al.
Veröffentlicht: (2025)
von: Lian, Shijie, et al.
Veröffentlicht: (2025)
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
von: Li, Hongxing, et al.
Veröffentlicht: (2025)
von: Li, Hongxing, et al.
Veröffentlicht: (2025)
VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events
von: Bhat, Mohammad Qazim, et al.
Veröffentlicht: (2026)
von: Bhat, Mohammad Qazim, et al.
Veröffentlicht: (2026)
DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving
von: Fu, Yongjie, et al.
Veröffentlicht: (2024)
von: Fu, Yongjie, et al.
Veröffentlicht: (2024)
Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
von: Guo, Mingning, et al.
Veröffentlicht: (2025)
von: Guo, Mingning, et al.
Veröffentlicht: (2025)
CARE Drive A Framework for Evaluating Reason-Responsiveness of Vision Language Models in Automated Driving
von: Suryana, Lucas Elbert, et al.
Veröffentlicht: (2026)
von: Suryana, Lucas Elbert, et al.
Veröffentlicht: (2026)
Adversarial Prompt Distillation for Vision-Language Models
von: Luo, Lin, et al.
Veröffentlicht: (2024)
von: Luo, Lin, et al.
Veröffentlicht: (2024)
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024)
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
von: Song, Python, et al.
Veröffentlicht: (2025)
von: Song, Python, et al.
Veröffentlicht: (2025)
Evolving Prompt Adaptation for Vision-Language Models
von: Zhang, Enming, et al.
Veröffentlicht: (2026)
von: Zhang, Enming, et al.
Veröffentlicht: (2026)
LVLM-Aided Alignment of Task-Specific Vision Models
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)
SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving
von: Zheng, Peiru, et al.
Veröffentlicht: (2024)
von: Zheng, Peiru, et al.
Veröffentlicht: (2024)
RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving
von: Wang, Yujin, et al.
Veröffentlicht: (2025)
von: Wang, Yujin, et al.
Veröffentlicht: (2025)
RAC3: Retrieval-Augmented Corner Case Comprehension for Autonomous Driving with Vision-Language Models
von: Wang, Yujin, et al.
Veröffentlicht: (2024)
von: Wang, Yujin, et al.
Veröffentlicht: (2024)
Mask-RadarNet: Enhancing Transformer With Spatial-Temporal Semantic Context for Radar Object Detection in Autonomous Driving
von: Wu, Yuzhi, et al.
Veröffentlicht: (2024)
von: Wu, Yuzhi, et al.
Veröffentlicht: (2024)
Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent
von: He, Linfeng, et al.
Veröffentlicht: (2024)
von: He, Linfeng, et al.
Veröffentlicht: (2024)
Task-Specific Adaptation of Segmentation Foundation Model via Prompt Learning
von: Kim, Hyung-Il, et al.
Veröffentlicht: (2024)
von: Kim, Hyung-Il, et al.
Veröffentlicht: (2024)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
von: Singh, Ayush, et al.
Veröffentlicht: (2024)
von: Singh, Ayush, et al.
Veröffentlicht: (2024)
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
von: Yang, Sheng, et al.
Veröffentlicht: (2025)
von: Yang, Sheng, et al.
Veröffentlicht: (2025)
DriveCritic: Towards Context-Aware, Human-Aligned Evaluation for Autonomous Driving with Vision-Language Models
von: Song, Jingyu, et al.
Veröffentlicht: (2025)
von: Song, Jingyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
von: Tian, Kexin, et al.
Veröffentlicht: (2025) -
SpaceSense-Bench: A Large-Scale Multi-Modal Benchmark for Spacecraft Perception and Pose Estimation
von: Wu, Aodi, et al.
Veröffentlicht: (2026) -
Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
von: Yu, Seungjun, et al.
Veröffentlicht: (2025) -
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
von: Mei, Xiaodong, et al.
Veröffentlicht: (2026) -
INSIGHT: Enhancing Autonomous Driving Safety through Vision-Language Models on Context-Aware Hazard Detection and Edge Case Evaluation
von: Chen, Dianwei, et al.
Veröffentlicht: (2025)