RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Boyang, Zhang, Haoran, Zhang, Shujie, Hao, Jinkun, Jia, Mingda, Lv, Qi, Mao, Yucheng, Lyu, Zhaoyang, Zeng, Jia, Xu, Xudong, Pang, Jiangmiao
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915717083824128
author Wang, Boyang
Zhang, Haoran
Zhang, Shujie
Hao, Jinkun
Jia, Mingda
Lv, Qi
Mao, Yucheng
Lyu, Zhaoyang
Zeng, Jia
Xu, Xudong
Pang, Jiangmiao
author_facet Wang, Boyang
Zhang, Haoran
Zhang, Shujie
Hao, Jinkun
Jia, Mingda
Lv, Qi
Mao, Yucheng
Lyu, Zhaoyang
Zeng, Jia
Xu, Xudong
Pang, Jiangmiao
contents The diversity, quantity, and quality of manipulation data are critical for training effective robot policies. However, due to hardware and physical setup constraints, collecting large-scale real-world manipulation data remains difficult to scale across diverse environments. Recent work uses text-prompt conditioned image diffusion models to augment manipulation data by altering the backgrounds and tabletop objects in the visual observations. However, these approaches often overlook the practical need for multi-view and temporally coherent observations required by state-of-the-art policy models. Further, text prompts alone cannot reliably specify the scene setup. To provide the diffusion model with explicit visual guidance, we introduce visual identity prompting, which supplies exemplar images as conditioning inputs to guide the generation of the desired scene setup. To this end, we also build a scalable pipeline to curate a visual identity pool from large robotics datasets. Using our augmented manipulation data to train downstream vision-language-action and visuomotor policy models yields consistent performance gains in both simulation and real-robot settings.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05241
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
Wang, Boyang
Zhang, Haoran
Zhang, Shujie
Hao, Jinkun
Jia, Mingda
Lv, Qi
Mao, Yucheng
Lyu, Zhaoyang
Zeng, Jia
Xu, Xudong
Pang, Jiangmiao
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
The diversity, quantity, and quality of manipulation data are critical for training effective robot policies. However, due to hardware and physical setup constraints, collecting large-scale real-world manipulation data remains difficult to scale across diverse environments. Recent work uses text-prompt conditioned image diffusion models to augment manipulation data by altering the backgrounds and tabletop objects in the visual observations. However, these approaches often overlook the practical need for multi-view and temporally coherent observations required by state-of-the-art policy models. Further, text prompts alone cannot reliably specify the scene setup. To provide the diffusion model with explicit visual guidance, we introduce visual identity prompting, which supplies exemplar images as conditioning inputs to guide the generation of the desired scene setup. To this end, we also build a scalable pipeline to curate a visual identity pool from large robotics datasets. Using our augmented manipulation data to train downstream vision-language-action and visuomotor policy models yields consistent performance gains in both simulation and real-robot settings.
title RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2601.05241