Is an object-centric representation beneficial for robotic manipulation ?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chapin, Alexandre, Dellandrea, Emmanuel, Chen, Liming
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909658861535232
author Chapin, Alexandre
Dellandrea, Emmanuel
Chen, Liming
author_facet Chapin, Alexandre
Dellandrea, Emmanuel
Chen, Liming
contents Object-centric representation (OCR) has recently become a subject of interest in the computer vision community for learning a structured representation of images and videos. It has been several times presented as a potential way to improve data-efficiency and generalization capabilities to learn an agent on downstream tasks. However, most existing work only evaluates such models on scene decomposition, without any notion of reasoning over the learned representation. Robotic manipulation tasks generally involve multi-object environments with potential inter-object interaction. We thus argue that they are a very interesting playground to really evaluate the potential of existing object-centric work. To do so, we create several robotic manipulation tasks in simulated environments involving multiple objects (several distractors, the robot, etc.) and a high-level of randomization (object positions, colors, shapes, background, initial positions, etc.). We then evaluate one classical object-centric method across several generalization scenarios and compare its results against several state-of-the-art hollistic representations. Our results exhibit that existing methods are prone to failure in difficult scenarios involving complex scene structures, whereas object-centric methods help overcome these challenges.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19408
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Is an object-centric representation beneficial for robotic manipulation ?
Chapin, Alexandre
Dellandrea, Emmanuel
Chen, Liming
Artificial Intelligence
Robotics
Object-centric representation (OCR) has recently become a subject of interest in the computer vision community for learning a structured representation of images and videos. It has been several times presented as a potential way to improve data-efficiency and generalization capabilities to learn an agent on downstream tasks. However, most existing work only evaluates such models on scene decomposition, without any notion of reasoning over the learned representation. Robotic manipulation tasks generally involve multi-object environments with potential inter-object interaction. We thus argue that they are a very interesting playground to really evaluate the potential of existing object-centric work. To do so, we create several robotic manipulation tasks in simulated environments involving multiple objects (several distractors, the robot, etc.) and a high-level of randomization (object positions, colors, shapes, background, initial positions, etc.). We then evaluate one classical object-centric method across several generalization scenarios and compare its results against several state-of-the-art hollistic representations. Our results exhibit that existing methods are prone to failure in difficult scenarios involving complex scene structures, whereas object-centric methods help overcome these challenges.
title Is an object-centric representation beneficial for robotic manipulation ?
topic Artificial Intelligence
Robotics
url https://arxiv.org/abs/2506.19408