PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Xiaogang, Wang, Qian, Wang, Anrui, Wang, Han A., Gyenes, Balázs, Gospodinov, Emiliyan, Jiang, Xinkai, Li, Ge, Zhou, Hongyi, Liao, Weiran, Huang, Xi, Beck, Maximilian, Reuss, Moritz, Lioutikov, Rudolf, Neumann, Gerhard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911395119890432
author Jia, Xiaogang
Wang, Qian
Wang, Anrui
Wang, Han A.
Gyenes, Balázs
Gospodinov, Emiliyan
Jiang, Xinkai
Li, Ge
Zhou, Hongyi
Liao, Weiran
Huang, Xi
Beck, Maximilian
Reuss, Moritz
Lioutikov, Rudolf
Neumann, Gerhard
author_facet Jia, Xiaogang
Wang, Qian
Wang, Anrui
Wang, Han A.
Gyenes, Balázs
Gospodinov, Emiliyan
Jiang, Xinkai
Li, Ge
Zhou, Hongyi
Liao, Weiran
Huang, Xi
Beck, Maximilian
Reuss, Moritz
Lioutikov, Rudolf
Neumann, Gerhard
contents Robotic manipulation systems benefit from complementary sensing modalities, where each provides unique environmental information. Point clouds capture detailed geometric structure, while RGB images provide rich semantic context. Current point cloud methods struggle to capture fine-grained detail, especially for complex tasks, which RGB methods lack geometric awareness, which hinders their precision and generalization. We introduce PointMapPolicy, a novel approach that conditions diffusion policies on structured grids of points without downsampling. The resulting data type makes it easier to extract shape and spatial relationships from observations, and can be transformed between reference frames. Yet due to their structure in a regular grid, we enable the use of established computer vision techniques directly to 3D data. Using xLSTM as a backbone, our model efficiently fuses the point maps with RGB data for enhanced multi-modal perception. Through extensive experiments on the RoboCasa and CALVIN benchmarks and real robot evaluations, we demonstrate that our method achieves state-of-the-art performance across diverse manipulation tasks. The overview and demos are available on our project page: https://point-map.github.io/Point-Map/
format Preprint
id arxiv_https___arxiv_org_abs_2510_20406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning
Jia, Xiaogang
Wang, Qian
Wang, Anrui
Wang, Han A.
Gyenes, Balázs
Gospodinov, Emiliyan
Jiang, Xinkai
Li, Ge
Zhou, Hongyi
Liao, Weiran
Huang, Xi
Beck, Maximilian
Reuss, Moritz
Lioutikov, Rudolf
Neumann, Gerhard
Robotics
Machine Learning
Robotic manipulation systems benefit from complementary sensing modalities, where each provides unique environmental information. Point clouds capture detailed geometric structure, while RGB images provide rich semantic context. Current point cloud methods struggle to capture fine-grained detail, especially for complex tasks, which RGB methods lack geometric awareness, which hinders their precision and generalization. We introduce PointMapPolicy, a novel approach that conditions diffusion policies on structured grids of points without downsampling. The resulting data type makes it easier to extract shape and spatial relationships from observations, and can be transformed between reference frames. Yet due to their structure in a regular grid, we enable the use of established computer vision techniques directly to 3D data. Using xLSTM as a backbone, our model efficiently fuses the point maps with RGB data for enhanced multi-modal perception. Through extensive experiments on the RoboCasa and CALVIN benchmarks and real robot evaluations, we demonstrate that our method achieves state-of-the-art performance across diverse manipulation tasks. The overview and demos are available on our project page: https://point-map.github.io/Point-Map/
title PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning
topic Robotics
Machine Learning
url https://arxiv.org/abs/2510.20406