A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Liangdong, Nie, Yiming, Li, Haoyang, Kong, Fanjie, Zhang, Baobao, Huang, Shunxin, Fu, Kai, Min, Chen, Xiao, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks
by: Min, Chen, et al.
Published: (2025)
by: Min, Chen, et al.
Published: (2025)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025)
by: Hu, Tianshuai, et al.
Published: (2025)
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
by: Min, Chen, et al.
Published: (2023)
by: Min, Chen, et al.
Published: (2023)
StyleVLA: Driving Style-Aware Vision Language Action Model for Autonomous Driving
by: Gao, Yuan, et al.
Published: (2026)
by: Gao, Yuan, et al.
Published: (2026)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
Autonomous Driving in Unstructured Environments: How Far Have We Come?
by: Min, Chen, et al.
Published: (2024)
by: Min, Chen, et al.
Published: (2024)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
Wild-Drive: Off-Road Scene Captioning and Path Planning via Robust Multi-modal Routing and Efficient Large Language Model
by: Wang, Zihang, et al.
Published: (2026)
by: Wang, Zihang, et al.
Published: (2026)
Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
by: Zhang, Yichi, et al.
Published: (2026)
by: Zhang, Yichi, et al.
Published: (2026)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
by: Huang, Wenhui, et al.
Published: (2026)
by: Huang, Wenhui, et al.
Published: (2026)
Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models
by: Jia, Xiaosong, et al.
Published: (2026)
by: Jia, Xiaosong, et al.
Published: (2026)
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
by: Zhang, Dapeng, et al.
Published: (2025)
by: Zhang, Dapeng, et al.
Published: (2025)
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
by: Huang, Helong, et al.
Published: (2025)
by: Huang, Helong, et al.
Published: (2025)
Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent
by: He, Linfeng, et al.
Published: (2024)
by: He, Linfeng, et al.
Published: (2024)
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
by: Chen, Guanyan, et al.
Published: (2024)
by: Chen, Guanyan, et al.
Published: (2024)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model
by: Zhang, Dapeng, et al.
Published: (2025)
by: Zhang, Dapeng, et al.
Published: (2025)
Seeing before Observable: Potential Risk Reasoning in Autonomous Driving via Vision Language Models
by: Liu, Jiaxin, et al.
Published: (2025)
by: Liu, Jiaxin, et al.
Published: (2025)
Dynamics Modeling using Visual Terrain Features for High-Speed Autonomous Off-Road Driving
by: Gibson, Jason, et al.
Published: (2024)
by: Gibson, Jason, et al.
Published: (2024)
DeeAD: Dynamic Early Exit of Vision-Language Action for Efficient Autonomous Driving
by: HU, Haibo, et al.
Published: (2025)
by: HU, Haibo, et al.
Published: (2025)
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
by: Wu, Yanhao, et al.
Published: (2026)
by: Wu, Yanhao, et al.
Published: (2026)
SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
by: Renz, Katrin, et al.
Published: (2025)
by: Renz, Katrin, et al.
Published: (2025)
Using Language and Road Manuals to Inform Map Reconstruction for Autonomous Driving
by: Tumu, Akshar, et al.
Published: (2025)
by: Tumu, Akshar, et al.
Published: (2025)
Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
by: Lee, Sangoh, et al.
Published: (2025)
by: Lee, Sangoh, et al.
Published: (2025)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
AppleVLM: End-to-end Autonomous Driving with Advanced Perception and Planning-Enhanced Vision-Language Models
by: Han, Yuxuan, et al.
Published: (2026)
by: Han, Yuxuan, et al.
Published: (2026)
CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine
by: Fang, Shiyu, et al.
Published: (2025)
by: Fang, Shiyu, et al.
Published: (2025)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
by: Zhang, Zezhou, et al.
Published: (2026)
by: Zhang, Zezhou, et al.
Published: (2026)
Benchmarking and Improving Bird's Eye View Perception Robustness in Autonomous Driving
by: Xie, Shaoyuan, et al.
Published: (2024)
by: Xie, Shaoyuan, et al.
Published: (2024)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
by: Zhong, Yifan, et al.
Published: (2025)
by: Zhong, Yifan, et al.
Published: (2025)
MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
by: Fu, Yuxia, et al.
Published: (2025)
by: Fu, Yuxia, et al.
Published: (2025)
Self-Supervised Traversability Learning with Online Prototype Adaptation for Off-Road Autonomous Driving
by: Bu, Yafeng, et al.
Published: (2025)
by: Bu, Yafeng, et al.
Published: (2025)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
by: Li, Pengteng, et al.
Published: (2026)
by: Li, Pengteng, et al.
Published: (2026)
DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
by: Yang, Zhenjie, et al.
Published: (2025)
by: Yang, Zhenjie, et al.
Published: (2025)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
by: Liang, Yuanchang, et al.
Published: (2026)
by: Liang, Yuanchang, et al.
Published: (2026)
Causal Scene Narration with Runtime Safety Supervision for Vision-Language-Action Driving
by: Li, Yun, et al.
Published: (2026)
by: Li, Yun, et al.
Published: (2026)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Similar Items
-
Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks
by: Min, Chen, et al.
Published: (2025) -
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025) -
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
by: Fu, Haoyu, et al.
Published: (2025) -
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025) -
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
by: Min, Chen, et al.
Published: (2023)