VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913058039791616 |
|---|---|
| author | Gu, Songen Zheng, Yuhang Li, Weize Zheng, Yupeng Feng, Yating Li, Xiang Chen, Yilun Li, Pengfei Ding, Wenchao |
| author_facet | Gu, Songen Zheng, Yuhang Li, Weize Zheng, Yupeng Feng, Yating Li, Xiang Chen, Yilun Li, Pengfei Ding, Wenchao |
| contents | Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera. In this paper, we propose VistaBot, a novel framework that integrates feed-forward geometric models with video diffusion models to achieve view-robust closed-loop manipulation without requiring camera calibration at test time. Our approach consists of three key components: 4D geometry estimation, view synthesis latent extraction, and latent action learning. VistaBot is integrated into both action-chunking (ACT) and diffusion-based ($π_0$) policies and evaluated across simulation and real-world tasks. We further introduce the View Generalization Score (VGS) as a new metric for comprehensive evaluation of cross-view generalization. Results show that VistaBot improves VGS by 2.79$\times$ and 2.63$\times$ over ACT and $π_0$, respectively, while also achieving high-quality novel view synthesis. Our contributions include a geometry-aware synthesis model, a latent action planner, a new benchmark metric, and extensive validation across diverse environments. The code and models will be made publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_21914 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis Gu, Songen Zheng, Yuhang Li, Weize Zheng, Yupeng Feng, Yating Li, Xiang Chen, Yilun Li, Pengfei Ding, Wenchao Robotics Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera. In this paper, we propose VistaBot, a novel framework that integrates feed-forward geometric models with video diffusion models to achieve view-robust closed-loop manipulation without requiring camera calibration at test time. Our approach consists of three key components: 4D geometry estimation, view synthesis latent extraction, and latent action learning. VistaBot is integrated into both action-chunking (ACT) and diffusion-based ($π_0$) policies and evaluated across simulation and real-world tasks. We further introduce the View Generalization Score (VGS) as a new metric for comprehensive evaluation of cross-view generalization. Results show that VistaBot improves VGS by 2.79$\times$ and 2.63$\times$ over ACT and $π_0$, respectively, while also achieving high-quality novel view synthesis. Our contributions include a geometry-aware synthesis model, a latent action planner, a new benchmark metric, and extensive validation across diverse environments. The code and models will be made publicly available. |
| title | VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis |
| topic | Robotics |
| url | https://arxiv.org/abs/2604.21914 |