Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917484429312000 |
|---|---|
| author | Xiao, Junjin Li, Dongyang Yang, Yandan Zeng, Shuang Lin, Tong Chang, Xinyuan Xiong, Feng Xu, Mu Wei, Xing Ma, Zhiheng Zhang, Qing Zheng, Wei-Shi |
| author_facet | Xiao, Junjin Li, Dongyang Yang, Yandan Zeng, Shuang Lin, Tong Chang, Xinyuan Xiong, Feng Xu, Mu Wei, Xing Ma, Zhiheng Zhang, Qing Zheng, Wei-Shi |
| contents | This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained multi-view diffusion model to synthesize latent novel views and propose a Geometry-Guided Gated Transformer (G3T) that aligns multi-view features under 3D geometric guidance while adaptively filtering occlusion noise. To improve action learning efficiency, we introduce Action Manifold Learning (AML), which directly predicts actions on the valid action manifold, bypassing inefficient regression of unstructured targets like noise or velocity. Experiments on LIBERO, RoboTwin 2.0, and real-robot tasks show our method achieves superior success rate and robustness over SOTA baselines. Project page: https://junjxiao.github.io/Multi-view-VLA.github.io/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_11832 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Xiao, Junjin Li, Dongyang Yang, Yandan Zeng, Shuang Lin, Tong Chang, Xinyuan Xiong, Feng Xu, Mu Wei, Xing Ma, Zhiheng Zhang, Qing Zheng, Wei-Shi Robotics This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained multi-view diffusion model to synthesize latent novel views and propose a Geometry-Guided Gated Transformer (G3T) that aligns multi-view features under 3D geometric guidance while adaptively filtering occlusion noise. To improve action learning efficiency, we introduce Action Manifold Learning (AML), which directly predicts actions on the valid action manifold, bypassing inefficient regression of unstructured targets like noise or velocity. Experiments on LIBERO, RoboTwin 2.0, and real-robot tasks show our method achieves superior success rate and robustness over SOTA baselines. Project page: https://junjxiao.github.io/Multi-view-VLA.github.io/. |
| title | Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation |
| topic | Robotics |
| url | https://arxiv.org/abs/2605.11832 |