VIP: Vision Instructed Pre-training for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhuoling, Ren, Liangliang, Yang, Jinrong, Zhao, Yong, Wu, Xiaoyang, Xu, Zhenhua, Bai, Xiang, Zhao, Hengshuang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sim-to-Real Dynamic Object Manipulation on Conveyor Systems via Optimization Path Shaping
by: Li, Zhuoling, et al.
Published: (2025)
by: Li, Zhuoling, et al.
Published: (2025)
Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space
by: Yang, Jinrong, et al.
Published: (2025)
by: Yang, Jinrong, et al.
Published: (2025)
ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot Learning
by: Jin, Zhao, et al.
Published: (2025)
by: Jin, Zhao, et al.
Published: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
InsMapper: Exploring Inner-instance Information for Vectorized HD Mapping
by: Xu, Zhenhua, et al.
Published: (2023)
by: Xu, Zhenhua, et al.
Published: (2023)
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
by: Fan, Xianzhe, et al.
Published: (2026)
by: Fan, Xianzhe, et al.
Published: (2026)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
by: Jiang, Guangqi, et al.
Published: (2024)
by: Jiang, Guangqi, et al.
Published: (2024)
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
by: Wang, Boyang, et al.
Published: (2026)
by: Wang, Boyang, et al.
Published: (2026)
PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation
by: Yin, Yifan, et al.
Published: (2025)
by: Yin, Yifan, et al.
Published: (2025)
Spatiotemporal Predictive Pre-training for Robotic Motor Control
by: Yang, Jiange, et al.
Published: (2024)
by: Yang, Jiange, et al.
Published: (2024)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
by: Xie, Senwei, et al.
Published: (2025)
by: Xie, Senwei, et al.
Published: (2025)
Confusion-Aware In-Context-Learning for Vision-Language Models in Robotic Manipulation
by: He, Yayun, et al.
Published: (2026)
by: He, Yayun, et al.
Published: (2026)
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
Task-oriented Robotic Manipulation with Vision Language Models
by: Guran, Nurhan Bulus, et al.
Published: (2024)
by: Guran, Nurhan Bulus, et al.
Published: (2024)
SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
by: Wu, Shengkai, et al.
Published: (2025)
by: Wu, Shengkai, et al.
Published: (2025)
VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation
by: Zhao, Wentao, et al.
Published: (2024)
by: Zhao, Wentao, et al.
Published: (2024)
CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
by: Song, Kun, et al.
Published: (2025)
by: Song, Kun, et al.
Published: (2025)
DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
by: Xu, Zhenhua, et al.
Published: (2023)
by: Xu, Zhenhua, et al.
Published: (2023)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
by: Fan, Yiguo, et al.
Published: (2025)
by: Fan, Yiguo, et al.
Published: (2025)
RoboRouter: Training-Free Policy Routing for Robotic Manipulation
by: Chen, Yiteng, et al.
Published: (2026)
by: Chen, Yiteng, et al.
Published: (2026)
T-araVLN: Translator for Agricultural Robotic Agents on Vision-and-Language Navigation
by: Zhao, Xiaobei, et al.
Published: (2025)
by: Zhao, Xiaobei, et al.
Published: (2025)
Closed-Loop Magnetic Manipulation for Robotic Transesophageal Echocardiography
by: Li, Keyu, et al.
Published: (2023)
by: Li, Keyu, et al.
Published: (2023)
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
GhostObjects: Instructing Robots by Manipulating Spatially Aligned Virtual Twins in Augmented Reality
by: Wang, Lauren W., et al.
Published: (2025)
by: Wang, Lauren W., et al.
Published: (2025)
SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model
by: Zhu, Shaoting, et al.
Published: (2024)
by: Zhu, Shaoting, et al.
Published: (2024)
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
by: Zuo, Kuangji, et al.
Published: (2026)
by: Zuo, Kuangji, et al.
Published: (2026)
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
by: Bai, Shuanghao, et al.
Published: (2025)
by: Bai, Shuanghao, et al.
Published: (2025)
VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving
by: Zhang, Haiming, et al.
Published: (2024)
by: Zhang, Haiming, et al.
Published: (2024)
Instruct2Act: From Human Instruction to Actions Sequencing and Execution via Robot Action Network for Robotic Manipulation
by: Sharma, Archit, et al.
Published: (2026)
by: Sharma, Archit, et al.
Published: (2026)
DML-RAM: Deep Multimodal Learning Framework for Robotic Arm Manipulation using Pre-trained Models
by: Kumar, Sathish, et al.
Published: (2025)
by: Kumar, Sathish, et al.
Published: (2025)
When would Vision-Proprioception Policies Fail in Robotic Manipulation?
by: Lu, Jingxian, et al.
Published: (2026)
by: Lu, Jingxian, et al.
Published: (2026)
A Novel Planning Framework for Complex Flipping Manipulation of Multiple Mobile Manipulators
by: Liu, Wenhang, et al.
Published: (2023)
by: Liu, Wenhang, et al.
Published: (2023)
SimLauncher: Launching Sample-Efficient Real-world Robotic Reinforcement Learning via Simulation Pre-training
by: Wu, Mingdong, et al.
Published: (2025)
by: Wu, Mingdong, et al.
Published: (2025)
Learning Shared RGB-D Fields: Unified Self-supervised Pre-training for Label-efficient LiDAR-Camera 3D Perception
by: Xu, Xiaohao, et al.
Published: (2024)
by: Xu, Xiaohao, et al.
Published: (2024)
Lifelike Agility and Play in Quadrupedal Robots using Reinforcement Learning and Generative Pre-trained Models
by: Han, Lei, et al.
Published: (2023)
by: Han, Lei, et al.
Published: (2023)
Towards Generalizable Robotic Manipulation in Dynamic Environments
by: Fang, Heng, et al.
Published: (2026)
by: Fang, Heng, et al.
Published: (2026)
VADF: Vision-Adaptive Diffusion Policy Framework for Efficient Robotic Manipulation
by: Yu, Xinglei, et al.
Published: (2026)
by: Yu, Xinglei, et al.
Published: (2026)
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Triple Regression for Camera Agnostic Sim2Real Robot Grasping and Manipulation Tasks
by: Zeng, Yuanhong, et al.
Published: (2023)
by: Zeng, Yuanhong, et al.
Published: (2023)
Similar Items
-
Sim-to-Real Dynamic Object Manipulation on Conveyor Systems via Optimization Path Shaping
by: Li, Zhuoling, et al.
Published: (2025) -
Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space
by: Yang, Jinrong, et al.
Published: (2025) -
ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot Learning
by: Jin, Zhao, et al.
Published: (2025) -
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025) -
InsMapper: Exploring Inner-instance Information for Vectorized HD Mapping
by: Xu, Zhenhua, et al.
Published: (2023)