DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Son, Dongwon, Son, Sanghyeon, Kim, Jaehyung, Kim, Beomjoon
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916341089304576
author Son, Dongwon
Son, Sanghyeon
Kim, Jaehyung
Kim, Beomjoon
author_facet Son, Dongwon
Son, Sanghyeon
Kim, Jaehyung
Kim, Beomjoon
contents We present DEF-oriCORN, a framework for language-directed manipulation tasks. By leveraging a novel object-based scene representation and diffusion-model-based state estimation algorithm, our framework enables efficient and robust manipulation planning in response to verbal commands, even in tightly packed environments with sparse camera views without any demonstrations. Unlike traditional representations, our representation affords efficient collision checking and language grounding. Compared to state-of-the-art baselines, our framework achieves superior estimation and motion planning performance from sparse RGB images and zero-shot generalizes to real-world scenarios with diverse materials, including transparent and reflective objects, despite being trained exclusively in simulation. Our code for data generation, training, inference, and pre-trained weights are publicly available at: https://sites.google.com/view/def-oricorn/home.
format Preprint
id arxiv_https___arxiv_org_abs_2407_21267
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
Son, Dongwon
Son, Sanghyeon
Kim, Jaehyung
Kim, Beomjoon
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
We present DEF-oriCORN, a framework for language-directed manipulation tasks. By leveraging a novel object-based scene representation and diffusion-model-based state estimation algorithm, our framework enables efficient and robust manipulation planning in response to verbal commands, even in tightly packed environments with sparse camera views without any demonstrations. Unlike traditional representations, our representation affords efficient collision checking and language grounding. Compared to state-of-the-art baselines, our framework achieves superior estimation and motion planning performance from sparse RGB images and zero-shot generalizes to real-world scenarios with diverse materials, including transparent and reflective objects, despite being trained exclusively in simulation. Our code for data generation, training, inference, and pre-trained weights are publicly available at: https://sites.google.com/view/def-oricorn/home.
title DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.21267