FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Xiaoxu, Li, Hao, Ye, Jinhui, Chen, Yilun, Zeng, Jia, Chen, Xinyi, Xu, Linning, Lin, Dahua, Li, Weixin, Pang, Jiangmiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
by: Yang, Sizhe, et al.
Published: (2026)
by: Yang, Sizhe, et al.
Published: (2026)
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
by: Tian, Yang, et al.
Published: (2024)
by: Tian, Yang, et al.
Published: (2024)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
by: Zheng, Jinliang, et al.
Published: (2025)
by: Zheng, Jinliang, et al.
Published: (2025)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
by: Chen, Xinyi, et al.
Published: (2025)
by: Chen, Xinyi, et al.
Published: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
by: Chi, Cheng, et al.
Published: (2023)
by: Chi, Cheng, et al.
Published: (2023)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
Long-Term Memory for VLA-based Agents in Open-World Task Execution
by: Huang, Xu, et al.
Published: (2026)
by: Huang, Xu, et al.
Published: (2026)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
by: Li, Boyu, et al.
Published: (2026)
by: Li, Boyu, et al.
Published: (2026)
DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos
by: Mu, Juncheng, et al.
Published: (2026)
by: Mu, Juncheng, et al.
Published: (2026)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
by: Liang, Zhixuan, et al.
Published: (2025)
by: Liang, Zhixuan, et al.
Published: (2025)
HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit
by: Ben, Qingwei, et al.
Published: (2025)
by: Ben, Qingwei, et al.
Published: (2025)
UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data
by: Yang, Sizhe, et al.
Published: (2026)
by: Yang, Sizhe, et al.
Published: (2026)
GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
by: Gao, Ning, et al.
Published: (2025)
by: Gao, Ning, et al.
Published: (2025)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
by: Zheng, Yupeng, et al.
Published: (2026)
by: Zheng, Yupeng, et al.
Published: (2026)
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
by: Li, Xiaoqi, et al.
Published: (2026)
by: Li, Xiaoqi, et al.
Published: (2026)
RynnVLA-002: A Unified Vision-Language-Action and World Model
by: Cen, Jun, et al.
Published: (2025)
by: Cen, Jun, et al.
Published: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
by: Li, Yixuan, et al.
Published: (2025)
by: Li, Yixuan, et al.
Published: (2025)
Learning Predictive Visuomotor Coordination
by: Jia, Wenqi, et al.
Published: (2025)
by: Jia, Wenqi, et al.
Published: (2025)
One-Policy-Fits-All: Geometry-Aware Action Latents for Cross-Embodiment Manipulation
by: Mu, Juncheng, et al.
Published: (2026)
by: Mu, Juncheng, et al.
Published: (2026)
Learning H-Infinity Locomotion Control
by: Long, Junfeng, et al.
Published: (2024)
by: Long, Junfeng, et al.
Published: (2024)
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
by: Qu, Delin, et al.
Published: (2025)
by: Qu, Delin, et al.
Published: (2025)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
Gallant: Voxel Grid-based Humanoid Locomotion and Local-navigation across 3D Constrained Terrains
by: Ben, Qingwei, et al.
Published: (2025)
by: Ben, Qingwei, et al.
Published: (2025)
RationalVLA: A Rational Vision-Language-Action Model with Dual System
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
by: Zhong, Yifan, et al.
Published: (2025)
by: Zhong, Yifan, et al.
Published: (2025)
AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation
by: Zhu, Juan, et al.
Published: (2026)
by: Zhu, Juan, et al.
Published: (2026)
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
by: Cai, Junhao, et al.
Published: (2026)
by: Cai, Junhao, et al.
Published: (2026)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
by: Peng, Xiongfeng, et al.
Published: (2026)
by: Peng, Xiongfeng, et al.
Published: (2026)
Similar Items
-
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
by: Li, Hao, et al.
Published: (2025) -
Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
by: Yang, Sizhe, et al.
Published: (2026) -
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
by: Ye, Jinhui, et al.
Published: (2026) -
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025) -
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)