MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiaxu, Jiang, Yicheng, He, Tianlun, Sun, Jingkai, Zhang, Qiang, He, Junhao, Cao, Jiahang, Gan, Zesen, Sun, Mingyuan, Shao, Qiming, Yue, Xiangyu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918524427960320
author Wang, Jiaxu
Jiang, Yicheng
He, Tianlun
Sun, Jingkai
Zhang, Qiang
He, Junhao
Cao, Jiahang
Gan, Zesen
Sun, Mingyuan
Shao, Qiming
Yue, Xiangyu
author_facet Wang, Jiaxu
Jiang, Yicheng
He, Tianlun
Sun, Jingkai
Zhang, Qiang
He, Junhao
Cao, Jiahang
Gan, Zesen
Sun, Mingyuan
Shao, Qiming
Yue, Xiangyu
contents World-model-based imagine-then-act becomes a promising paradigm for robotic manipulation, yet existing approaches typically support either purely image-based forecasting or reasoning over partial 3D geometry, limiting their ability to predict complete 4D scene dynamics. This work proposes a novel embodied 4D world model that enables geometrically consistent, arbitrary-view RGBD generation: given only a single-view RGBD observation as input, the model imagines the remaining viewpoints, which can then be back-projected and fused to assemble a more complete 3D structure across time. To efficiently learn the multi-view, cross-modality generation, we explicitly design cross-view and cross-modality feature fusion that jointly encourage consistency between RGB and depth and enforce geometric alignment across views. Beyond prediction, converting generated futures into actions is often handled by inverse dynamics, which is ill-posed because multiple actions can explain the same transition. We address this with a test-time action optimization strategy that backpropagates through the generative model to infer a trajectory-level latent best matching the predicted future, and a residual inverse dynamics model that turns this trajectory prior into accurate executable actions. Experiments on three datasets demonstrate strong performance on both 4D scene generation and downstream manipulation, and ablations provide practical insights into the key design choices.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09878
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation
Wang, Jiaxu
Jiang, Yicheng
He, Tianlun
Sun, Jingkai
Zhang, Qiang
He, Junhao
Cao, Jiahang
Gan, Zesen
Sun, Mingyuan
Shao, Qiming
Yue, Xiangyu
Computer Vision and Pattern Recognition
World-model-based imagine-then-act becomes a promising paradigm for robotic manipulation, yet existing approaches typically support either purely image-based forecasting or reasoning over partial 3D geometry, limiting their ability to predict complete 4D scene dynamics. This work proposes a novel embodied 4D world model that enables geometrically consistent, arbitrary-view RGBD generation: given only a single-view RGBD observation as input, the model imagines the remaining viewpoints, which can then be back-projected and fused to assemble a more complete 3D structure across time. To efficiently learn the multi-view, cross-modality generation, we explicitly design cross-view and cross-modality feature fusion that jointly encourage consistency between RGB and depth and enforce geometric alignment across views. Beyond prediction, converting generated futures into actions is often handled by inverse dynamics, which is ill-posed because multiple actions can explain the same transition. We address this with a test-time action optimization strategy that backpropagates through the generative model to infer a trajectory-level latent best matching the predicted future, and a residual inverse dynamics model that turns this trajectory prior into accurate executable actions. Experiments on three datasets demonstrate strong performance on both 4D scene generation and downstream manipulation, and ablations provide practical insights into the key design choices.
title MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.09878