VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Seong, Hyunki, Moon, Seongwoo, Ahn, Hojin, Kang, Jehun, Shim, David Hyunchul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Supervised Interpretable End-to-End Learning via Latent Functional Modularity
by: Seong, Hyunki, et al.
Published: (2024)
by: Seong, Hyunki, et al.
Published: (2024)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
MonoDINO-DETR: Depth-Enhanced Monocular 3D Object Detection Using a Vision Foundation Model
by: Kim, Jihyeok, et al.
Published: (2025)
by: Kim, Jihyeok, et al.
Published: (2025)
Words to Wheels: Vision-Based Autonomous Driving Understanding Human Language Instructions Using Foundation Models
by: Ryu, Chanhoe, et al.
Published: (2024)
by: Ryu, Chanhoe, et al.
Published: (2024)
ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving
by: Sheng, Zihao, et al.
Published: (2026)
by: Sheng, Zihao, et al.
Published: (2026)
AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning
by: Zhou, Zewei, et al.
Published: (2025)
by: Zhou, Zewei, et al.
Published: (2025)
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
by: Wang, Yiru, et al.
Published: (2026)
by: Wang, Yiru, et al.
Published: (2026)
Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion
by: Zhang, Jiaru, et al.
Published: (2026)
by: Zhang, Jiaru, et al.
Published: (2026)
SUPER-AD: Semantic Uncertainty-aware Planning for End-to-End Robust Autonomous Driving
by: Ryu, Wonjeong, et al.
Published: (2025)
by: Ryu, Wonjeong, et al.
Published: (2025)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026)
by: Wang, Linbo, et al.
Published: (2026)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
by: Yang, Zhenjie, et al.
Published: (2025)
by: Yang, Zhenjie, et al.
Published: (2025)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
by: Chen, Xuesong, et al.
Published: (2025)
by: Chen, Xuesong, et al.
Published: (2025)
MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving
by: Sun, Bin, et al.
Published: (2025)
by: Sun, Bin, et al.
Published: (2025)
Multimodal Action Diffusion for Robust End-to-End Autonomous Driving
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving
by: Song, Nan, et al.
Published: (2025)
by: Song, Nan, et al.
Published: (2025)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
Enhancing End-to-End Autonomous Driving with Latent World Model
by: Li, Yingyan, et al.
Published: (2024)
by: Li, Yingyan, et al.
Published: (2024)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving
by: Arai, Hidehisa, et al.
Published: (2024)
by: Arai, Hidehisa, et al.
Published: (2024)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
by: Huang, Wenhui, et al.
Published: (2026)
by: Huang, Wenhui, et al.
Published: (2026)
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
by: Zhang, Songyan, et al.
Published: (2024)
by: Zhang, Songyan, et al.
Published: (2024)
ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving
by: Zhang, Jinqing, et al.
Published: (2026)
by: Zhang, Jinqing, et al.
Published: (2026)
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
by: Han, Jianhua, et al.
Published: (2025)
by: Han, Jianhua, et al.
Published: (2025)
HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models
by: Cho, Hoonhee, et al.
Published: (2026)
by: Cho, Hoonhee, et al.
Published: (2026)
E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
by: Tang, Yihong, et al.
Published: (2025)
by: Tang, Yihong, et al.
Published: (2025)
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
by: Chi, Haohan, et al.
Published: (2025)
by: Chi, Haohan, et al.
Published: (2025)
TempFuser: Learning Agile, Tactical, and Acrobatic Flight Maneuvers Using a Long Short-Term Temporal Fusion Transformer
by: Seong, Hyunki, et al.
Published: (2023)
by: Seong, Hyunki, et al.
Published: (2023)
Skill Q-Network: Learning Adaptive Skill Ensemble for Mapless Navigation in Unknown Environments
by: Seong, Hyunki, et al.
Published: (2024)
by: Seong, Hyunki, et al.
Published: (2024)
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
by: Jiang, Hao, et al.
Published: (2025)
by: Jiang, Hao, et al.
Published: (2025)
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
by: Zheng, Yupeng, et al.
Published: (2025)
by: Zheng, Yupeng, et al.
Published: (2025)
OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
by: Zhang, Songyan, et al.
Published: (2025)
by: Zhang, Songyan, et al.
Published: (2025)
AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
by: Xu, Peng, et al.
Published: (2026)
by: Xu, Peng, et al.
Published: (2026)
AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models
by: Yan, Tianyi, et al.
Published: (2025)
by: Yan, Tianyi, et al.
Published: (2025)
OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
by: Xing, Shuo, et al.
Published: (2024)
by: Xing, Shuo, et al.
Published: (2024)
ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving
by: Cui, Can, et al.
Published: (2025)
by: Cui, Can, et al.
Published: (2025)
VECTOR-Drive: Tightly Coupled Vision-Language and Trajectory Expert Routing for End-to-End Autonomous Driving
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
by: Chen, Yi, et al.
Published: (2026)
by: Chen, Yi, et al.
Published: (2026)
Similar Items
-
Self-Supervised Interpretable End-to-End Learning via Latent Functional Modularity
by: Seong, Hyunki, et al.
Published: (2024) -
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
by: Zhou, Xingcheng, et al.
Published: (2025) -
MonoDINO-DETR: Depth-Enhanced Monocular 3D Object Detection Using a Vision Foundation Model
by: Kim, Jihyeok, et al.
Published: (2025) -
Words to Wheels: Vision-Based Autonomous Driving Understanding Human Language Instructions Using Foundation Models
by: Ryu, Chanhoe, et al.
Published: (2024) -
ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving
by: Sheng, Zihao, et al.
Published: (2026)