Representing Positional Information in Generative World Models for Object Manipulation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ferraro, Stefano, Mazzaglia, Pietro, Verbelen, Tim, Dhoedt, Bart, Rajeswar, Sai
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912034613886976
author Ferraro, Stefano
Mazzaglia, Pietro
Verbelen, Tim
Dhoedt, Bart
Rajeswar, Sai
author_facet Ferraro, Stefano
Mazzaglia, Pietro
Verbelen, Tim
Dhoedt, Bart
Rajeswar, Sai
contents Object manipulation capabilities are essential skills that set apart embodied agents engaging with the world, especially in the realm of robotics. The ability to predict outcomes of interactions with objects is paramount in this setting. While model-based control methods have started to be employed for tackling manipulation tasks, they have faced challenges in accurately manipulating objects. As we analyze the causes of this limitation, we identify the cause of underperformance in the way current world models represent crucial positional information, especially about the target's goal specification for object positioning tasks. We introduce a general approach that empowers world model-based agents to effectively solve object-positioning tasks. We propose two declinations of this approach for generative world models: position-conditioned (PCP) and latent-conditioned (LCP) policy learning. In particular, LCP employs object-centric latent representations that explicitly capture object positional information for goal specification. This naturally leads to the emergence of multimodal capabilities, enabling the specification of goals through spatial coordinates or a visual goal. Our methods are rigorously evaluated across several manipulation environments, showing favorable performance compared to current model-based control approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2409_12005
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Representing Positional Information in Generative World Models for Object Manipulation
Ferraro, Stefano
Mazzaglia, Pietro
Verbelen, Tim
Dhoedt, Bart
Rajeswar, Sai
Robotics
Artificial Intelligence
Object manipulation capabilities are essential skills that set apart embodied agents engaging with the world, especially in the realm of robotics. The ability to predict outcomes of interactions with objects is paramount in this setting. While model-based control methods have started to be employed for tackling manipulation tasks, they have faced challenges in accurately manipulating objects. As we analyze the causes of this limitation, we identify the cause of underperformance in the way current world models represent crucial positional information, especially about the target's goal specification for object positioning tasks. We introduce a general approach that empowers world model-based agents to effectively solve object-positioning tasks. We propose two declinations of this approach for generative world models: position-conditioned (PCP) and latent-conditioned (LCP) policy learning. In particular, LCP employs object-centric latent representations that explicitly capture object positional information for goal specification. This naturally leads to the emergence of multimodal capabilities, enabling the specification of goals through spatial coordinates or a visual goal. Our methods are rigorously evaluated across several manipulation environments, showing favorable performance compared to current model-based control approaches.
title Representing Positional Information in Generative World Models for Object Manipulation
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2409.12005