VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Songen, Zheng, Yuhang, Li, Weize, Zheng, Yupeng, Feng, Yating, Li, Xiang, Chen, Yilun, Li, Pengfei, Ding, Wenchao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
by: Zheng, Yuhang, et al.
Published: (2026)
by: Zheng, Yuhang, et al.
Published: (2026)
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
by: Zheng, Yupeng, et al.
Published: (2026)
by: Zheng, Yupeng, et al.
Published: (2026)
World In Your Hands: A Large-Scale and Open-Source Ecosystem for Learning Human-Centric Manipulation in the Wild
by: Zheng, Yupeng, et al.
Published: (2025)
by: Zheng, Yupeng, et al.
Published: (2025)
Learning High-Frequency Continuous Action Chunks in Latent Space
by: Wang, Kunyun, et al.
Published: (2026)
by: Wang, Kunyun, et al.
Published: (2026)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
by: Liu, Fanfan, et al.
Published: (2024)
by: Liu, Fanfan, et al.
Published: (2024)
ForeDiffusion: Foresight-Conditioned Diffusion Policy via Future View Construction for Robot Manipulation
by: Xie, Weize, et al.
Published: (2026)
by: Xie, Weize, et al.
Published: (2026)
Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation
by: Gu, Songen, et al.
Published: (2026)
by: Gu, Songen, et al.
Published: (2026)
GaussianGrasper: 3D Language Gaussian Splatting for Open-vocabulary Robotic Grasping
by: Zheng, Yuhang, et al.
Published: (2024)
by: Zheng, Yuhang, et al.
Published: (2024)
Correlation-Aware Dual-View Pose and Velocity Estimation for Dynamic Robotic Manipulation
by: Zarei, Mahboubeh, et al.
Published: (2025)
by: Zarei, Mahboubeh, et al.
Published: (2025)
UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation
by: Jin, Bu, et al.
Published: (2025)
by: Jin, Bu, et al.
Published: (2025)
Cortical Policy: A Dual-Stream View Transformer for Robotic Manipulation
by: Zhang, Xuening, et al.
Published: (2026)
by: Zhang, Xuening, et al.
Published: (2026)
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
by: Bai, Yongjie, et al.
Published: (2025)
by: Bai, Yongjie, et al.
Published: (2025)
Language-Guided Object-Centric Diffusion Policy for Generalizable and Collision-Aware Robotic Manipulation
by: Li, Hang, et al.
Published: (2024)
by: Li, Hang, et al.
Published: (2024)
EEG-Driven AR-Robot System for Zero-Touch Grasping Manipulation
by: Wang, Junzhe, et al.
Published: (2025)
by: Wang, Junzhe, et al.
Published: (2025)
Learning Geometrically-Grounded 3D Visual Representations for View-Generalizable Robotic Manipulation
by: Zhang, Di, et al.
Published: (2026)
by: Zhang, Di, et al.
Published: (2026)
Data Scaling Laws for Imitation Learning-Based End-to-End Autonomous Driving
by: Zheng, Yupeng, et al.
Published: (2024)
by: Zheng, Yupeng, et al.
Published: (2024)
Mimir: Hierarchical Goal-Driven Diffusion with Uncertainty Propagation for End-to-End Autonomous Driving
by: Xing, Zebin, et al.
Published: (2025)
by: Xing, Zebin, et al.
Published: (2025)
ShakingBot: Dynamic Manipulation for Bagging
by: Gu, Ningquan, et al.
Published: (2023)
by: Gu, Ningquan, et al.
Published: (2023)
ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis
by: Fang, Yu, et al.
Published: (2025)
by: Fang, Yu, et al.
Published: (2025)
FlowBotHD: History-Aware Diffuser Handling Ambiguities in Articulated Objects Manipulation
by: Li, Yishu, et al.
Published: (2024)
by: Li, Yishu, et al.
Published: (2024)
Rhythm: Learning Interactive Whole-Body Control for Dual Humanoids
by: Chen, Hongjin, et al.
Published: (2026)
by: Chen, Hongjin, et al.
Published: (2026)
SAMP: Spatial Anchor-based Motion Policy for Collision-Aware Robotic Manipulators
by: Chen, Kai, et al.
Published: (2025)
by: Chen, Kai, et al.
Published: (2025)
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
by: Lan, Zihan, et al.
Published: (2025)
by: Lan, Zihan, et al.
Published: (2025)
GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
by: Qian, Quanhao, et al.
Published: (2025)
by: Qian, Quanhao, et al.
Published: (2025)
Enhanced View Planning for Robotic Harvesting: Tackling Occlusions with Imitation Learning
by: Li, Lun, et al.
Published: (2025)
by: Li, Lun, et al.
Published: (2025)
Beyond Viewpoint Generalization: What Multi-View Demonstrations Offer and How to Synthesize Them for Robot Manipulation?
by: Cai, Boyang, et al.
Published: (2026)
by: Cai, Boyang, et al.
Published: (2026)
ManiVID-3D: Generalizable View-Invariant Reinforcement Learning for Robotic Manipulation via Disentangled 3D Representations
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
Visual Robotic Manipulation with Depth-Aware Pretraining
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
M4Diffuser: Multi-View Diffusion Policy with Manipulability-Aware Control for Robust Mobile Manipulation
by: Dong, Ju, et al.
Published: (2025)
by: Dong, Ju, et al.
Published: (2025)
A0: An Affordance-Aware Hierarchical Model for General Robotic Manipulation
by: Xu, Rongtao, et al.
Published: (2025)
by: Xu, Rongtao, et al.
Published: (2025)
ST-VLA: Enabling 4D-Aware Spatiotemporal Understanding for General Robot Manipulation
by: Wu, You, et al.
Published: (2026)
by: Wu, You, et al.
Published: (2026)
$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation
by: Zhou, Pengfei, et al.
Published: (2026)
by: Zhou, Pengfei, et al.
Published: (2026)
Imitation Learning for Active Neck Motion Enabling Robot Manipulation beyond the Field of View
by: Nakagawa, Koki, et al.
Published: (2025)
by: Nakagawa, Koki, et al.
Published: (2025)
Enhancing Indoor Occupancy Prediction via Sparse Query-Based Multi-Level Consistent Knowledge Distillation
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
OminiAdapt: Learning Cross-Task Invariance for Robust and Environment-Aware Robotic Manipulation
by: Wang, Yongxu, et al.
Published: (2025)
by: Wang, Yongxu, et al.
Published: (2025)
Locomotion as Manipulation with ReachBot
by: Chen, Tony G., et al.
Published: (2024)
by: Chen, Tony G., et al.
Published: (2024)
HyperTASR: Hypernetwork-Driven Task-Aware Scene Representations for Robust Manipulation
by: Sun, Li, et al.
Published: (2025)
by: Sun, Li, et al.
Published: (2025)
WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
by: Qian, Zezhong, et al.
Published: (2025)
by: Qian, Zezhong, et al.
Published: (2025)
FACTO: Function-space Adaptive Constrained Trajectory Optimization for Robotic Manipulators
by: Feng, Yichang, et al.
Published: (2026)
by: Feng, Yichang, et al.
Published: (2026)
RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Similar Items
-
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
by: Zheng, Yuhang, et al.
Published: (2026) -
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
by: Zheng, Yupeng, et al.
Published: (2026) -
World In Your Hands: A Large-Scale and Open-Source Ecosystem for Learning Human-Centric Manipulation in the Wild
by: Zheng, Yupeng, et al.
Published: (2025) -
Learning High-Frequency Continuous Action Chunks in Latent Space
by: Wang, Kunyun, et al.
Published: (2026) -
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
by: Liu, Fanfan, et al.
Published: (2024)