VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Songen, Zheng, Yuhang, Li, Weize, Zheng, Yupeng, Feng, Yating, Li, Xiang, Chen, Yilun, Li, Pengfei, Ding, Wenchao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913058039791616
author Gu, Songen
Zheng, Yuhang
Li, Weize
Zheng, Yupeng
Feng, Yating
Li, Xiang
Chen, Yilun
Li, Pengfei
Ding, Wenchao
author_facet Gu, Songen
Zheng, Yuhang
Li, Weize
Zheng, Yupeng
Feng, Yating
Li, Xiang
Chen, Yilun
Li, Pengfei
Ding, Wenchao
contents Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera. In this paper, we propose VistaBot, a novel framework that integrates feed-forward geometric models with video diffusion models to achieve view-robust closed-loop manipulation without requiring camera calibration at test time. Our approach consists of three key components: 4D geometry estimation, view synthesis latent extraction, and latent action learning. VistaBot is integrated into both action-chunking (ACT) and diffusion-based ($π_0$) policies and evaluated across simulation and real-world tasks. We further introduce the View Generalization Score (VGS) as a new metric for comprehensive evaluation of cross-view generalization. Results show that VistaBot improves VGS by 2.79$\times$ and 2.63$\times$ over ACT and $π_0$, respectively, while also achieving high-quality novel view synthesis. Our contributions include a geometry-aware synthesis model, a latent action planner, a new benchmark metric, and extensive validation across diverse environments. The code and models will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21914
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
Gu, Songen
Zheng, Yuhang
Li, Weize
Zheng, Yupeng
Feng, Yating
Li, Xiang
Chen, Yilun
Li, Pengfei
Ding, Wenchao
Robotics
Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera. In this paper, we propose VistaBot, a novel framework that integrates feed-forward geometric models with video diffusion models to achieve view-robust closed-loop manipulation without requiring camera calibration at test time. Our approach consists of three key components: 4D geometry estimation, view synthesis latent extraction, and latent action learning. VistaBot is integrated into both action-chunking (ACT) and diffusion-based ($π_0$) policies and evaluated across simulation and real-world tasks. We further introduce the View Generalization Score (VGS) as a new metric for comprehensive evaluation of cross-view generalization. Results show that VistaBot improves VGS by 2.79$\times$ and 2.63$\times$ over ACT and $π_0$, respectively, while also achieving high-quality novel view synthesis. Our contributions include a geometry-aware synthesis model, a latent action planner, a new benchmark metric, and extensive validation across diverse environments. The code and models will be made publicly available.
title VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
topic Robotics
url https://arxiv.org/abs/2604.21914