PointPatchRL -- Masked Reconstruction Improves Reinforcement Learning on Point Clouds

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gyenes, Balázs, Franke, Nikolai, Becker, Philipp, Neumann, Gerhard
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914987262345216
author Gyenes, Balázs
Franke, Nikolai
Becker, Philipp
Neumann, Gerhard
author_facet Gyenes, Balázs
Franke, Nikolai
Becker, Philipp
Neumann, Gerhard
contents Perceiving the environment via cameras is crucial for Reinforcement Learning (RL) in robotics. While images are a convenient form of representation, they often complicate extracting important geometric details, especially with varying geometries or deformable objects. In contrast, point clouds naturally represent this geometry and easily integrate color and positional data from multiple camera views. However, while deep learning on point clouds has seen many recent successes, RL on point clouds is under-researched, with only the simplest encoder architecture considered in the literature. We introduce PointPatchRL (PPRL), a method for RL on point clouds that builds on the common paradigm of dividing point clouds into overlapping patches, tokenizing them, and processing the tokens with transformers. PPRL provides significant improvements compared with other point-cloud processing architectures previously used for RL. We then complement PPRL with masked reconstruction for representation learning and show that our method outperforms strong model-free and model-based baselines on image observations in complex manipulation tasks containing deformable objects and variations in target object geometry. Videos and code are available at https://alrhub.github.io/pprl-website
format Preprint
id arxiv_https___arxiv_org_abs_2410_18800
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PointPatchRL -- Masked Reconstruction Improves Reinforcement Learning on Point Clouds
Gyenes, Balázs
Franke, Nikolai
Becker, Philipp
Neumann, Gerhard
Machine Learning
Robotics
Perceiving the environment via cameras is crucial for Reinforcement Learning (RL) in robotics. While images are a convenient form of representation, they often complicate extracting important geometric details, especially with varying geometries or deformable objects. In contrast, point clouds naturally represent this geometry and easily integrate color and positional data from multiple camera views. However, while deep learning on point clouds has seen many recent successes, RL on point clouds is under-researched, with only the simplest encoder architecture considered in the literature. We introduce PointPatchRL (PPRL), a method for RL on point clouds that builds on the common paradigm of dividing point clouds into overlapping patches, tokenizing them, and processing the tokens with transformers. PPRL provides significant improvements compared with other point-cloud processing architectures previously used for RL. We then complement PPRL with masked reconstruction for representation learning and show that our method outperforms strong model-free and model-based baselines on image observations in complex manipulation tasks containing deformable objects and variations in target object geometry. Videos and code are available at https://alrhub.github.io/pprl-website
title PointPatchRL -- Masked Reconstruction Improves Reinforcement Learning on Point Clouds
topic Machine Learning
Robotics
url https://arxiv.org/abs/2410.18800