Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Junjin, Li, Dongyang, Yang, Yandan, Zeng, Shuang, Lin, Tong, Chang, Xinyuan, Xiong, Feng, Xu, Mu, Wei, Xing, Ma, Zhiheng, Zhang, Qing, Zheng, Wei-Shi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917484429312000
author Xiao, Junjin
Li, Dongyang
Yang, Yandan
Zeng, Shuang
Lin, Tong
Chang, Xinyuan
Xiong, Feng
Xu, Mu
Wei, Xing
Ma, Zhiheng
Zhang, Qing
Zheng, Wei-Shi
author_facet Xiao, Junjin
Li, Dongyang
Yang, Yandan
Zeng, Shuang
Lin, Tong
Chang, Xinyuan
Xiong, Feng
Xu, Mu
Wei, Xing
Ma, Zhiheng
Zhang, Qing
Zheng, Wei-Shi
contents This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained multi-view diffusion model to synthesize latent novel views and propose a Geometry-Guided Gated Transformer (G3T) that aligns multi-view features under 3D geometric guidance while adaptively filtering occlusion noise. To improve action learning efficiency, we introduce Action Manifold Learning (AML), which directly predicts actions on the valid action manifold, bypassing inefficient regression of unstructured targets like noise or velocity. Experiments on LIBERO, RoboTwin 2.0, and real-robot tasks show our method achieves superior success rate and robustness over SOTA baselines. Project page: https://junjxiao.github.io/Multi-view-VLA.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11832
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation
Xiao, Junjin
Li, Dongyang
Yang, Yandan
Zeng, Shuang
Lin, Tong
Chang, Xinyuan
Xiong, Feng
Xu, Mu
Wei, Xing
Ma, Zhiheng
Zhang, Qing
Zheng, Wei-Shi
Robotics
This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained multi-view diffusion model to synthesize latent novel views and propose a Geometry-Guided Gated Transformer (G3T) that aligns multi-view features under 3D geometric guidance while adaptively filtering occlusion noise. To improve action learning efficiency, we introduce Action Manifold Learning (AML), which directly predicts actions on the valid action manifold, bypassing inefficient regression of unstructured targets like noise or velocity. Experiments on LIBERO, RoboTwin 2.0, and real-robot tasks show our method achieves superior success rate and robustness over SOTA baselines. Project page: https://junjxiao.github.io/Multi-view-VLA.github.io/.
title Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation
topic Robotics
url https://arxiv.org/abs/2605.11832