What Matters in Building Vision-Language-Action Models for Generalist Robots
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinghang, Li, Peiyan, Qian, Long, Liu, Minghuan, Wang, Dong, Liu, Jirong, Kang, Bingyi, Ma, Xiao, Wang, Xinlong, Guo, Di, Kong, Tao, Zhang, Hanbo, Liu, Huaping |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-Language Foundation Models as Effective Robot Imitators
by: Li, Xinghang, et al.
Published: (2023)
by: Li, Xinghang, et al.
Published: (2023)
SInViG: A Self-Evolving Interactive Visual Agent for Human-Robot Interaction
by: Xu, Jie, et al.
Published: (2024)
by: Xu, Jie, et al.
Published: (2024)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
by: Wang, Guokang, et al.
Published: (2024)
by: Wang, Guokang, et al.
Published: (2024)
Scaling World Model for Hierarchical Manipulation Policies
by: Long, Qian, et al.
Published: (2026)
by: Long, Qian, et al.
Published: (2026)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
by: Guo, Jun, et al.
Published: (2026)
by: Guo, Jun, et al.
Published: (2026)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
Demonstrating HumanTHOR: A Simulation Platform and Benchmark for Human-Robot Collaboration in a Shared Workspace
by: Wang, Chenxu, et al.
Published: (2024)
by: Wang, Chenxu, et al.
Published: (2024)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
by: Liu, Minghuan, et al.
Published: (2025)
by: Liu, Minghuan, et al.
Published: (2025)
Leveraging Large Language Model for Heterogeneous Ad Hoc Teamwork Collaboration
by: Liu, Xinzhu, et al.
Published: (2024)
by: Liu, Xinzhu, et al.
Published: (2024)
Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
by: Li, Peiyan, et al.
Published: (2026)
by: Li, Peiyan, et al.
Published: (2026)
CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
by: Sun, Nan, et al.
Published: (2025)
by: Sun, Nan, et al.
Published: (2025)
Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
by: Liu, Huihan, et al.
Published: (2026)
by: Liu, Huihan, et al.
Published: (2026)
Sparse Spectral Training and Inference on Euclidean and Hyperbolic Neural Networks
by: Zhao, Jialin, et al.
Published: (2024)
by: Zhao, Jialin, et al.
Published: (2024)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
by: Zhang, Zhengshen, et al.
Published: (2025)
by: Zhang, Zhengshen, et al.
Published: (2025)
Vision Generalist Model: A Survey
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
by: Apanasevich, I., et al.
Published: (2026)
by: Apanasevich, I., et al.
Published: (2026)
Data Readout Techniques for DNA‐Based Information Storage
by: Bingyi Liu, et al.
Published: (2025)
by: Bingyi Liu, et al.
Published: (2025)
Towards Objectively Benchmarking Social Intelligence for Language Agents at Action Level
by: Wang, Chenxu, et al.
Published: (2024)
by: Wang, Chenxu, et al.
Published: (2024)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
by: Chen, Xinyi, et al.
Published: (2025)
by: Chen, Xinyi, et al.
Published: (2025)
D4W: Dependable Data-Driven Dynamics for Wheeled Robots
by: Lin, Yunfeng, et al.
Published: (2024)
by: Lin, Yunfeng, et al.
Published: (2024)
Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
by: Sun, Nan, et al.
Published: (2025)
by: Sun, Nan, et al.
Published: (2025)
UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models
by: Yang, Jiabing, et al.
Published: (2026)
by: Yang, Jiabing, et al.
Published: (2026)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
by: Cheang, Chi-Lam, et al.
Published: (2024)
by: Cheang, Chi-Lam, et al.
Published: (2024)
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
by: Shi, Junyao, et al.
Published: (2025)
by: Shi, Junyao, et al.
Published: (2025)
Effective Tuning Strategies for Generalist Robot Manipulation Policies
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
by: Reuss, Moritz, et al.
Published: (2025)
by: Reuss, Moritz, et al.
Published: (2025)
NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
by: Chen, Jiahong, et al.
Published: (2025)
by: Chen, Jiahong, et al.
Published: (2025)
We Need Two Worlds'
by: Minghuan, Li
Published: (2010)
by: Minghuan, Li
Published: (2010)
Towards Robust Deep Reinforcement Learning against Environmental State Perturbation
by: Wang, Chenxu, et al.
Published: (2025)
by: Wang, Chenxu, et al.
Published: (2025)
Learning What Matters: Adaptive Information-Theoretic Objectives for Robot Exploration
by: Yu, Youwei, et al.
Published: (2026)
by: Yu, Youwei, et al.
Published: (2026)
MADiff: Offline Multi-agent Learning with Diffusion Models
by: Zhu, Zhengbang, et al.
Published: (2023)
by: Zhu, Zhengbang, et al.
Published: (2023)
Stimulate the Potential of Robots via Competition
by: Huang, Kangyao, et al.
Published: (2024)
by: Huang, Kangyao, et al.
Published: (2024)
Generalist Robot Manipulation beyond Action Labeled Data
by: Spiridonov, Alexander, et al.
Published: (2025)
by: Spiridonov, Alexander, et al.
Published: (2025)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
UltraMedical: Building Specialized Generalists in Biomedicine
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow
by: Chen, Yixiang, et al.
Published: (2025)
by: Chen, Yixiang, et al.
Published: (2025)
Temporal Action Localization with Cross Layer Task Decoupling and Refinement
by: Li, Qiang, et al.
Published: (2024)
by: Li, Qiang, et al.
Published: (2024)
Similar Items
-
Vision-Language Foundation Models as Effective Robot Imitators
by: Li, Xinghang, et al.
Published: (2023) -
SInViG: A Self-Evolving Interactive Visual Agent for Human-Robot Interaction
by: Xu, Jie, et al.
Published: (2024) -
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
by: Wang, Guokang, et al.
Published: (2024) -
Scaling World Model for Hierarchical Manipulation Policies
by: Long, Qian, et al.
Published: (2026) -
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
by: Guo, Jun, et al.
Published: (2026)