VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Haiming, Zhou, Wending, Zhu, Yiyao, Yan, Xu, Gao, Jiantao, Bai, Dongfeng, Cai, Yingjie, Liu, Bingbing, Cui, Shuguang, Li, Zhen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous Driving
by: Zhang, Haiming, et al.
Published: (2025)
by: Zhang, Haiming, et al.
Published: (2025)
DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving
by: Zhu, Yiyao, et al.
Published: (2026)
by: Zhu, Yiyao, et al.
Published: (2026)
Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models
by: Xu, Tianshuo, et al.
Published: (2025)
by: Xu, Tianshuo, et al.
Published: (2025)
UniPAD: A Universal Pre-training Paradigm for Autonomous Driving
by: Yang, Honghui, et al.
Published: (2023)
by: Yang, Honghui, et al.
Published: (2023)
DINO Pre-training for Vision-based End-to-end Autonomous Driving
by: Juneja, Shubham, et al.
Published: (2024)
by: Juneja, Shubham, et al.
Published: (2024)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025)
by: Hu, Tianshuai, et al.
Published: (2025)
VIP: Vision Instructed Pre-training for Robotic Manipulation
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
Forging Vision Foundation Models for Autonomous Driving: Challenges, Methodologies, and Opportunities
by: Yan, Xu, et al.
Published: (2024)
by: Yan, Xu, et al.
Published: (2024)
DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models
by: Zhong, Zhide, et al.
Published: (2026)
by: Zhong, Zhide, et al.
Published: (2026)
HUGSIM: A Real-Time, Photo-Realistic and Closed-Loop Simulator for Autonomous Driving
by: Zhou, Hongyu, et al.
Published: (2024)
by: Zhou, Hongyu, et al.
Published: (2024)
Prioritizing Perception-Guided Self-Supervision: A New Paradigm for Causal Modeling in End-to-End Autonomous Driving
by: Huang, Yi, et al.
Published: (2025)
by: Huang, Yi, et al.
Published: (2025)
StyleVLA: Driving Style-Aware Vision Language Action Model for Autonomous Driving
by: Gao, Yuan, et al.
Published: (2026)
by: Gao, Yuan, et al.
Published: (2026)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
by: Schäfer, Finn Rasmus, et al.
Published: (2026)
by: Schäfer, Finn Rasmus, et al.
Published: (2026)
Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models
by: Jia, Xiaosong, et al.
Published: (2026)
by: Jia, Xiaosong, et al.
Published: (2026)
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
Integration of Computer Vision with Adaptive Control for Autonomous Driving Using ADORE
by: Ahammed, Abu Shad, et al.
Published: (2025)
by: Ahammed, Abu Shad, et al.
Published: (2025)
CAPS: Context-Aware Priority Sampling for Enhanced Imitation Learning in Autonomous Driving
by: Mirkhani, Hamidreza, et al.
Published: (2025)
by: Mirkhani, Hamidreza, et al.
Published: (2025)
Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
by: Hu, Haibo, et al.
Published: (2025)
by: Hu, Haibo, et al.
Published: (2025)
VECTOR-Drive: Tightly Coupled Vision-Language and Trajectory Expert Routing for End-to-End Autonomous Driving
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
by: Yang, Zhenjie, et al.
Published: (2025)
by: Yang, Zhenjie, et al.
Published: (2025)
Seeing before Observable: Potential Risk Reasoning in Autonomous Driving via Vision Language Models
by: Liu, Jiaxin, et al.
Published: (2025)
by: Liu, Jiaxin, et al.
Published: (2025)
GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving
by: Ljungbergh, William, et al.
Published: (2025)
by: Ljungbergh, William, et al.
Published: (2025)
UnderwaterVLA: Dual-brain Vision-Language-Action architecture for Autonomous Underwater Navigation
by: Wang, Zhangyuan, et al.
Published: (2025)
by: Wang, Zhangyuan, et al.
Published: (2025)
ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models
by: Li, Puhao, et al.
Published: (2025)
by: Li, Puhao, et al.
Published: (2025)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
A Hierarchical Test Platform for Vision Language Model (VLM)-Integrated Real-World Autonomous Driving
by: Zhou, Yupeng, et al.
Published: (2025)
by: Zhou, Yupeng, et al.
Published: (2025)
A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving
by: Zhang, Liangdong, et al.
Published: (2026)
by: Zhang, Liangdong, et al.
Published: (2026)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
by: Xing, Shuo, et al.
Published: (2024)
by: Xing, Shuo, et al.
Published: (2024)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
by: Nie, Chang, et al.
Published: (2026)
by: Nie, Chang, et al.
Published: (2026)
NDST: Neural Driving Style Transfer for Human-Like Vision-Based Autonomous Driving
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
Vision-based DRL Autonomous Driving Agent with Sim2Real Transfer
by: Li, Dianzhao, et al.
Published: (2023)
by: Li, Dianzhao, et al.
Published: (2023)
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
by: Jiang, Anqing, et al.
Published: (2025)
by: Jiang, Anqing, et al.
Published: (2025)
Poutine: Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving
by: Rowe, Luke, et al.
Published: (2025)
by: Rowe, Luke, et al.
Published: (2025)
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
Fringe Projection Based Vision Pipeline for Autonomous Hard Drive Disassembly
by: Balasubramaniam, Badrinath, et al.
Published: (2026)
by: Balasubramaniam, Badrinath, et al.
Published: (2026)
Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
Words to Wheels: Vision-Based Autonomous Driving Understanding Human Language Instructions Using Foundation Models
by: Ryu, Chanhoe, et al.
Published: (2024)
by: Ryu, Chanhoe, et al.
Published: (2024)
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
by: Gao, Haoxiang, et al.
Published: (2025)
by: Gao, Haoxiang, et al.
Published: (2025)
From Words to Wheels: Automated Style-Customized Policy Generation for Autonomous Driving
by: Han, Xu, et al.
Published: (2024)
by: Han, Xu, et al.
Published: (2024)
Similar Items
-
SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous Driving
by: Zhang, Haiming, et al.
Published: (2025) -
DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving
by: Zhu, Yiyao, et al.
Published: (2026) -
Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models
by: Xu, Tianshuo, et al.
Published: (2025) -
UniPAD: A Universal Pre-training Paradigm for Autonomous Driving
by: Yang, Honghui, et al.
Published: (2023) -
DINO Pre-training for Vision-based End-to-end Autonomous Driving
by: Juneja, Shubham, et al.
Published: (2024)