Vision-Language Foundation Models as Effective Robot Imitators
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinghang, Liu, Minghuan, Zhang, Hanbo, Yu, Cunjun, Xu, Jie, Wu, Hongtao, Cheang, Chilam, Jing, Ya, Zhang, Weinan, Liu, Huaping, Li, Hang, Kong, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SInViG: A Self-Evolving Interactive Visual Agent for Human-Robot Interaction
by: Xu, Jie, et al.
Published: (2024)
by: Xu, Jie, et al.
Published: (2024)
What Matters in Building Vision-Language-Action Models for Generalist Robots
by: Li, Xinghang, et al.
Published: (2024)
by: Li, Xinghang, et al.
Published: (2024)
IRASim: A Fine-Grained World Model for Robot Manipulation
by: Zhu, Fangqi, et al.
Published: (2024)
by: Zhu, Fangqi, et al.
Published: (2024)
GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy
by: Li, Peiyan, et al.
Published: (2024)
by: Li, Peiyan, et al.
Published: (2024)
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
by: Cheang, Chi-Lam, et al.
Published: (2024)
by: Cheang, Chi-Lam, et al.
Published: (2024)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
by: Wang, Guokang, et al.
Published: (2024)
by: Wang, Guokang, et al.
Published: (2024)
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
by: Liu, Minghuan, et al.
Published: (2025)
by: Liu, Minghuan, et al.
Published: (2025)
INVIGORATE: Interactive Visual Grounding and Grasping in Clutter
by: Zhang, Hanbo, et al.
Published: (2021)
by: Zhang, Hanbo, et al.
Published: (2021)
Towards Unified Interactive Visual Grounding in The Wild
by: Xu, Jie, et al.
Published: (2024)
by: Xu, Jie, et al.
Published: (2024)
A Unified and General Humanoid Whole-Body Controller for Versatile Locomotion
by: Xue, Yufei, et al.
Published: (2025)
by: Xue, Yufei, et al.
Published: (2025)
Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
by: Sun, Nan, et al.
Published: (2025)
by: Sun, Nan, et al.
Published: (2025)
CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
by: Sun, Nan, et al.
Published: (2025)
by: Sun, Nan, et al.
Published: (2025)
D4W: Dependable Data-Driven Dynamics for Wheeled Robots
by: Lin, Yunfeng, et al.
Published: (2024)
by: Lin, Yunfeng, et al.
Published: (2024)
TeachingBot: Robot Teacher for Human Handwriting
by: Hou, Zhimin, et al.
Published: (2023)
by: Hou, Zhimin, et al.
Published: (2023)
Scaling World Model for Hierarchical Manipulation Policies
by: Long, Qian, et al.
Published: (2026)
by: Long, Qian, et al.
Published: (2026)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
World Model-based Perception for Visual Legged Locomotion
by: Lai, Hang, et al.
Published: (2024)
by: Lai, Hang, et al.
Published: (2024)
Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation
by: Han, Xiaoshen, et al.
Published: (2025)
by: Han, Xiaoshen, et al.
Published: (2025)
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
by: Liu, Ruixun, et al.
Published: (2025)
by: Liu, Ruixun, et al.
Published: (2025)
H-Zero: Cross-Humanoid Locomotion Pretraining Enables Few-shot Novel Embodiment Transfer
by: Lin, Yunfeng, et al.
Published: (2025)
by: Lin, Yunfeng, et al.
Published: (2025)
Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
by: Liu, Huihan, et al.
Published: (2026)
by: Liu, Huihan, et al.
Published: (2026)
Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
by: Zhang, Zhengshen, et al.
Published: (2025)
by: Zhang, Zhengshen, et al.
Published: (2025)
GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs
by: Hua, Pu, et al.
Published: (2024)
by: Hua, Pu, et al.
Published: (2024)
CANINE: Coaching Visually Impaired Users for Interactive Navigation with a Robot Guide Dog
by: Yu, Cunjun, et al.
Published: (2026)
by: Yu, Cunjun, et al.
Published: (2026)
What Foundation Models can Bring for Robot Learning in Manipulation : A Survey
by: Li, Dingzhe, et al.
Published: (2024)
by: Li, Dingzhe, et al.
Published: (2024)
SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model
by: Zhu, Shaoting, et al.
Published: (2024)
by: Zhu, Shaoting, et al.
Published: (2024)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
by: Yin, Zhenhan, et al.
Published: (2025)
by: Yin, Zhenhan, et al.
Published: (2025)
FABG : End-to-end Imitation Learning for Embodied Affective Human-Robot Interaction
by: Zhang, Yanghai, et al.
Published: (2025)
by: Zhang, Yanghai, et al.
Published: (2025)
Leveraging Large Language Model for Heterogeneous Ad Hoc Teamwork Collaboration
by: Liu, Xinzhu, et al.
Published: (2024)
by: Liu, Xinzhu, et al.
Published: (2024)
Stimulate the Potential of Robots via Competition
by: Huang, Kangyao, et al.
Published: (2024)
by: Huang, Kangyao, et al.
Published: (2024)
Beyond the Majority: Long-tail Imitation Learning for Robotic Manipulation
by: Zhu, Junhong, et al.
Published: (2026)
by: Zhu, Junhong, et al.
Published: (2026)
In-situ Value-aligned Human-Robot Interactions with Physical Constraints
by: Li, Hongtao, et al.
Published: (2025)
by: Li, Hongtao, et al.
Published: (2025)
Diff-Muscle: Efficient Learning for Musculoskeletal Robotic Table Tennis
by: Zhao, Wentao, et al.
Published: (2026)
by: Zhao, Wentao, et al.
Published: (2026)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
by: Guo, Jun, et al.
Published: (2026)
by: Guo, Jun, et al.
Published: (2026)
GR-3 Technical Report
by: Cheang, Chilam, et al.
Published: (2025)
by: Cheang, Chilam, et al.
Published: (2025)
LoopSR: Looping Sim-and-Real for Lifelong Policy Adaptation of Legged Robots
by: Wu, Peilin, et al.
Published: (2024)
by: Wu, Peilin, et al.
Published: (2024)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
EvoGymCM: Harnessing Continuous Material Stiffness for Soft Robot Co-Design
by: Shen, Le, et al.
Published: (2026)
by: Shen, Le, et al.
Published: (2026)
Robot Operation of Home Appliances by Reading User Manuals
by: Zhang, Jian, et al.
Published: (2025)
by: Zhang, Jian, et al.
Published: (2025)
Similar Items
-
SInViG: A Self-Evolving Interactive Visual Agent for Human-Robot Interaction
by: Xu, Jie, et al.
Published: (2024) -
What Matters in Building Vision-Language-Action Models for Generalist Robots
by: Li, Xinghang, et al.
Published: (2024) -
IRASim: A Fine-Grained World Model for Robot Manipulation
by: Zhu, Fangqi, et al.
Published: (2024) -
GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy
by: Li, Peiyan, et al.
Published: (2024) -
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
by: Cheang, Chi-Lam, et al.
Published: (2024)