A comparison of visual representations for real-world reinforcement learning in the context of vacuum gripping

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sutter, Nico, Hartmann, Valentin N., Coros, Stelian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910857712107520
author Sutter, Nico
Hartmann, Valentin N.
Coros, Stelian
author_facet Sutter, Nico
Hartmann, Valentin N.
Coros, Stelian
contents When manipulating objects in the real world, we need reactive feedback policies that take into account sensor information to inform decisions. This study aims to determine how different encoders can be used in a reinforcement learning (RL) framework to interpret the spatial environment in the local surroundings of a robot arm. Our investigation focuses on comparing real-world vision with 3D scene inputs, exploring new architectures in the process. We built on the SERL framework, providing us with a sample efficient and stable RL foundation we could build upon, while keeping training times minimal. The results of this study indicate that spatial information helps to significantly outperform the visual counterpart, tested on a box picking task with a vacuum gripper. The code and videos of the evaluations are available at https://github.com/nisutte/voxel-serl.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A comparison of visual representations for real-world reinforcement learning in the context of vacuum gripping
Sutter, Nico
Hartmann, Valentin N.
Coros, Stelian
Robotics
When manipulating objects in the real world, we need reactive feedback policies that take into account sensor information to inform decisions. This study aims to determine how different encoders can be used in a reinforcement learning (RL) framework to interpret the spatial environment in the local surroundings of a robot arm. Our investigation focuses on comparing real-world vision with 3D scene inputs, exploring new architectures in the process. We built on the SERL framework, providing us with a sample efficient and stable RL foundation we could build upon, while keeping training times minimal. The results of this study indicate that spatial information helps to significantly outperform the visual counterpart, tested on a box picking task with a vacuum gripper. The code and videos of the evaluations are available at https://github.com/nisutte/voxel-serl.
title A comparison of visual representations for real-world reinforcement learning in the context of vacuum gripping
topic Robotics
url https://arxiv.org/abs/2503.02405