VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jie, Li, Guang, Huang, Zhijian, Dang, Chenxu, Ye, Hangjun, Han, Yahong, Chen, Long |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DriveFine: Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving
von: Dang, Chenxu, et al.
Veröffentlicht: (2026)
von: Dang, Chenxu, et al.
Veröffentlicht: (2026)
SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning
von: Dang, Chenxu, et al.
Veröffentlicht: (2026)
von: Dang, Chenxu, et al.
Veröffentlicht: (2026)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
von: Mei, Xiaodong, et al.
Veröffentlicht: (2026)
von: Mei, Xiaodong, et al.
Veröffentlicht: (2026)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
von: You, Zihan, et al.
Veröffentlicht: (2026)
von: You, Zihan, et al.
Veröffentlicht: (2026)
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
von: Luo, Yuechen, et al.
Veröffentlicht: (2026)
von: Luo, Yuechen, et al.
Veröffentlicht: (2026)
ViSE: A Systematic Approach to Vision-Only Street-View Extrapolation
von: Tan, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Tan, Kaiyuan, et al.
Veröffentlicht: (2025)
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
von: Li, Fuhao, et al.
Veröffentlicht: (2025)
von: Li, Fuhao, et al.
Veröffentlicht: (2025)
DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
von: Diao, Muxi, et al.
Veröffentlicht: (2025)
von: Diao, Muxi, et al.
Veröffentlicht: (2025)
Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives
von: Wang, Junli, et al.
Veröffentlicht: (2026)
von: Wang, Junli, et al.
Veröffentlicht: (2026)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
von: Li, Yongkang, et al.
Veröffentlicht: (2026)
von: Li, Yongkang, et al.
Veröffentlicht: (2026)
PerlAD: Towards Enhanced Closed-loop End-to-end Autonomous Driving with Pseudo-simulation-based Reinforcement Learning
von: Gao, Yinfeng, et al.
Veröffentlicht: (2026)
von: Gao, Yinfeng, et al.
Veröffentlicht: (2026)
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
von: Li, Yongkang, et al.
Veröffentlicht: (2025)
von: Li, Yongkang, et al.
Veröffentlicht: (2025)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
von: Huang, Zhijian, et al.
Veröffentlicht: (2024)
von: Huang, Zhijian, et al.
Veröffentlicht: (2024)
MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving
von: Wang, Junli, et al.
Veröffentlicht: (2026)
von: Wang, Junli, et al.
Veröffentlicht: (2026)
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
von: Tian, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Tian, Xiaoyu, et al.
Veröffentlicht: (2024)
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
von: Fu, Haoyu, et al.
Veröffentlicht: (2025)
von: Fu, Haoyu, et al.
Veröffentlicht: (2025)
Spatial-aware Vision Language Model for Autonomous Driving
von: Wei, Weijie, et al.
Veröffentlicht: (2025)
von: Wei, Weijie, et al.
Veröffentlicht: (2025)
WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
von: Zhu, Ziyue, et al.
Veröffentlicht: (2025)
von: Zhu, Ziyue, et al.
Veröffentlicht: (2025)
Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving
von: Yang, Lijin, et al.
Veröffentlicht: (2026)
von: Yang, Lijin, et al.
Veröffentlicht: (2026)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
von: Tan, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Tan, Kaiyuan, et al.
Veröffentlicht: (2026)
Visual Adversarial Attack on Vision-Language Models for Autonomous Driving
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
Scale-Aware UAV-to-Satellite Cross-View Geo-Localization: A Semantic Geometric Approach
von: Ye, Yibin, et al.
Veröffentlicht: (2026)
von: Ye, Yibin, et al.
Veröffentlicht: (2026)
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
von: Zhang, Songyan, et al.
Veröffentlicht: (2024)
von: Zhang, Songyan, et al.
Veröffentlicht: (2024)
VLP: Vision Language Planning for Autonomous Driving
von: Pan, Chenbin, et al.
Veröffentlicht: (2024)
von: Pan, Chenbin, et al.
Veröffentlicht: (2024)
CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving
von: Zhang, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhang, Tianrui, et al.
Veröffentlicht: (2025)
A Survey of Vision Transformers in Autonomous Driving: Current Trends and Future Directions
von: Lai-Dang, Quoc-Vinh
Veröffentlicht: (2024)
von: Lai-Dang, Quoc-Vinh
Veröffentlicht: (2024)
EGM: Efficient Visual Grounding Language Models
von: Zhan, Guanqi, et al.
Veröffentlicht: (2026)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2026)
Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
von: Guo, Xiangyu, et al.
Veröffentlicht: (2025)
von: Guo, Xiangyu, et al.
Veröffentlicht: (2025)
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
von: Chen, Xuesong, et al.
Veröffentlicht: (2025)
von: Chen, Xuesong, et al.
Veröffentlicht: (2025)
A Survey on Vision-Language-Action Models for Autonomous Driving
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
von: Zhang, Ruifei, et al.
Veröffentlicht: (2025)
von: Zhang, Ruifei, et al.
Veröffentlicht: (2025)
Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving
von: Zhang, Enming, et al.
Veröffentlicht: (2025)
von: Zhang, Enming, et al.
Veröffentlicht: (2025)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
von: Huang, Wenhui, et al.
Veröffentlicht: (2026)
von: Huang, Wenhui, et al.
Veröffentlicht: (2026)
Towards Cross-View Point Correspondence in Vision-Language Models
von: Wang, Yipu, et al.
Veröffentlicht: (2025)
von: Wang, Yipu, et al.
Veröffentlicht: (2025)
Generalizing to Out-of-Sample Degradations via Model Reprogramming
von: Jiang, Runhua, et al.
Veröffentlicht: (2024)
von: Jiang, Runhua, et al.
Veröffentlicht: (2024)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
von: Jiang, Bo, et al.
Veröffentlicht: (2024)
von: Jiang, Bo, et al.
Veröffentlicht: (2024)
MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving
von: Zhang, Enming, et al.
Veröffentlicht: (2024)
von: Zhang, Enming, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DriveFine: Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving
von: Dang, Chenxu, et al.
Veröffentlicht: (2026) -
SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning
von: Dang, Chenxu, et al.
Veröffentlicht: (2026) -
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
von: Mei, Xiaodong, et al.
Veröffentlicht: (2026) -
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
von: jia, Feiyang, et al.
Veröffentlicht: (2026) -
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
von: You, Zihan, et al.
Veröffentlicht: (2026)