Out of Sight, Still in Mind: Reasoning and Planning about Unobserved Objects with Video Tracking Enabled Memory Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Yixuan, Yuan, Jialin, Kim, Chanho, Pradhan, Pupul, Chen, Bryan, Fuxin, Li, Hermans, Tucker
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911886461632512
author Huang, Yixuan
Yuan, Jialin
Kim, Chanho
Pradhan, Pupul
Chen, Bryan
Fuxin, Li
Hermans, Tucker
author_facet Huang, Yixuan
Yuan, Jialin
Kim, Chanho
Pradhan, Pupul
Chen, Bryan
Fuxin, Li
Hermans, Tucker
contents Robots need to have a memory of previously observed, but currently occluded objects to work reliably in realistic environments. We investigate the problem of encoding object-oriented memory into a multi-object manipulation reasoning and planning framework. We propose DOOM and LOOM, which leverage transformer relational dynamics to encode the history of trajectories given partial-view point clouds and an object discovery and tracking engine. Our approaches can perform multiple challenging tasks including reasoning with occluded objects, novel objects appearance, and object reappearance. Throughout our extensive simulation and real-world experiments, we find that our approaches perform well in terms of different numbers of objects and different numbers of distractor actions. Furthermore, we show our approaches outperform an implicit memory baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2309_15278
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Out of Sight, Still in Mind: Reasoning and Planning about Unobserved Objects with Video Tracking Enabled Memory Models
Huang, Yixuan
Yuan, Jialin
Kim, Chanho
Pradhan, Pupul
Chen, Bryan
Fuxin, Li
Hermans, Tucker
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Robots need to have a memory of previously observed, but currently occluded objects to work reliably in realistic environments. We investigate the problem of encoding object-oriented memory into a multi-object manipulation reasoning and planning framework. We propose DOOM and LOOM, which leverage transformer relational dynamics to encode the history of trajectories given partial-view point clouds and an object discovery and tracking engine. Our approaches can perform multiple challenging tasks including reasoning with occluded objects, novel objects appearance, and object reappearance. Throughout our extensive simulation and real-world experiments, we find that our approaches perform well in terms of different numbers of objects and different numbers of distractor actions. Furthermore, we show our approaches outperform an implicit memory baseline.
title Out of Sight, Still in Mind: Reasoning and Planning about Unobserved Objects with Video Tracking Enabled Memory Models
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2309.15278