DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yi, Ge, Yuying, Zhou, Hui, Ding, Mingyu, Ge, Yixiao, Liu, Xihui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024)
by: Chen, Yi, et al.
Published: (2024)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
by: Chen, Yi, et al.
Published: (2023)
by: Chen, Yi, et al.
Published: (2023)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026)
by: Wang, Linbo, et al.
Published: (2026)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
Latent Chain-of-Thought World Modeling for End-to-End Driving
by: Tan, Shuhan, et al.
Published: (2025)
by: Tan, Shuhan, et al.
Published: (2025)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
by: Sun, Jingwen, et al.
Published: (2026)
by: Sun, Jingwen, et al.
Published: (2026)
DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving
by: Zhang, Lingjun, et al.
Published: (2026)
by: Zhang, Lingjun, et al.
Published: (2026)
AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
by: Xu, Peng, et al.
Published: (2026)
by: Xu, Peng, et al.
Published: (2026)
World Guidance: World Modeling in Condition Space for Action Generation
by: Su, Yue, et al.
Published: (2026)
by: Su, Yue, et al.
Published: (2026)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
by: Qiu, Lu, et al.
Published: (2024)
by: Qiu, Lu, et al.
Published: (2024)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Unraveling the Effects of Synthetic Data on End-to-End Autonomous Driving
by: Ge, Junhao, et al.
Published: (2025)
by: Ge, Junhao, et al.
Published: (2025)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
by: Han, Jianhua, et al.
Published: (2025)
by: Han, Jianhua, et al.
Published: (2025)
DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
by: Fang, Zhen, et al.
Published: (2025)
by: Fang, Zhen, et al.
Published: (2025)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
by: Wang, Yating, et al.
Published: (2025)
by: Wang, Yating, et al.
Published: (2025)
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
by: Wang, Yiru, et al.
Published: (2026)
by: Wang, Yiru, et al.
Published: (2026)
VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
by: Li, Hengtao, et al.
Published: (2025)
by: Li, Hengtao, et al.
Published: (2025)
HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models
by: Cho, Hoonhee, et al.
Published: (2026)
by: Cho, Hoonhee, et al.
Published: (2026)
CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine
by: Fang, Shiyu, et al.
Published: (2025)
by: Fang, Shiyu, et al.
Published: (2025)
DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
by: Gao, Yu, et al.
Published: (2025)
by: Gao, Yu, et al.
Published: (2025)
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling
by: Lyu, Huaihai, et al.
Published: (2026)
by: Lyu, Huaihai, et al.
Published: (2026)
MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving
by: Yasarla, Rajeev, et al.
Published: (2026)
by: Yasarla, Rajeev, et al.
Published: (2026)
HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving
by: Yao, Wenhao, et al.
Published: (2026)
by: Yao, Wenhao, et al.
Published: (2026)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
by: Huang, Wenhui, et al.
Published: (2026)
by: Huang, Wenhui, et al.
Published: (2026)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
by: Zhang, Wenyao, et al.
Published: (2025)
by: Zhang, Wenyao, et al.
Published: (2025)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?
by: Xu, Yihong, et al.
Published: (2023)
by: Xu, Yihong, et al.
Published: (2023)
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
by: Zheng, Huan, et al.
Published: (2025)
by: Zheng, Huan, et al.
Published: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving
by: Shao, Hao, et al.
Published: (2026)
by: Shao, Hao, et al.
Published: (2026)
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
by: Qiu, Lu, et al.
Published: (2025)
by: Qiu, Lu, et al.
Published: (2025)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving
by: Shang, Shuyao, et al.
Published: (2026)
by: Shang, Shuyao, et al.
Published: (2026)
Latent Action Pretraining Through World Modeling
by: Tharwat, Bahey, et al.
Published: (2025)
by: Tharwat, Bahey, et al.
Published: (2025)
Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving
by: Sun, Jiangxin, et al.
Published: (2026)
by: Sun, Jiangxin, et al.
Published: (2026)
Collision Risk Estimation via Loss Prediction in End-to-End Autonomous Driving
by: Xiong, Ziliang, et al.
Published: (2025)
by: Xiong, Ziliang, et al.
Published: (2025)
Similar Items
-
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024) -
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
by: Chen, Yi, et al.
Published: (2023) -
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026) -
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025) -
Latent Chain-of-Thought World Modeling for End-to-End Driving
by: Tan, Shuhan, et al.
Published: (2025)