Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nguyen, Toan, Vu, Minh Nhat, Huang, Baoru, Vuong, An, Vuong, Quan, Le, Ngan, Vo, Thieu, Nguyen, Anh
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916335870541824
author Nguyen, Toan
Vu, Minh Nhat
Huang, Baoru
Vuong, An
Vuong, Quan
Le, Ngan
Vo, Thieu
Nguyen, Anh
author_facet Nguyen, Toan
Vu, Minh Nhat
Huang, Baoru
Vuong, An
Vuong, Quan
Le, Ngan
Vo, Thieu
Nguyen, Anh
contents 6-DoF grasp detection has been a fundamental and challenging problem in robotic vision. While previous works have focused on ensuring grasp stability, they often do not consider human intention conveyed through natural language, hindering effective collaboration between robots and users in complex 3D environments. In this paper, we present a new approach for language-driven 6-DoF grasp detection in cluttered point clouds. We first introduce Grasp-Anything-6D, a large-scale dataset for the language-driven 6-DoF grasp detection task with 1M point cloud scenes and more than 200M language-associated 3D grasp poses. We further introduce a novel diffusion model that incorporates a new negative prompt guidance learning strategy. The proposed negative prompt strategy directs the detection process toward the desired object while steering away from unwanted ones given the language input. Our method enables an end-to-end framework where humans can command the robot to grasp desired objects in a cluttered scene using natural language. Intensive experimental results show the effectiveness of our method in both benchmarking experiments and real-world scenarios, surpassing other baselines. In addition, we demonstrate the practicality of our approach in real-world robotic applications. Our project is available at https://airvlab.github.io/grasp-anything.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13842
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
Nguyen, Toan
Vu, Minh Nhat
Huang, Baoru
Vuong, An
Vuong, Quan
Le, Ngan
Vo, Thieu
Nguyen, Anh
Robotics
Computer Vision and Pattern Recognition
6-DoF grasp detection has been a fundamental and challenging problem in robotic vision. While previous works have focused on ensuring grasp stability, they often do not consider human intention conveyed through natural language, hindering effective collaboration between robots and users in complex 3D environments. In this paper, we present a new approach for language-driven 6-DoF grasp detection in cluttered point clouds. We first introduce Grasp-Anything-6D, a large-scale dataset for the language-driven 6-DoF grasp detection task with 1M point cloud scenes and more than 200M language-associated 3D grasp poses. We further introduce a novel diffusion model that incorporates a new negative prompt guidance learning strategy. The proposed negative prompt strategy directs the detection process toward the desired object while steering away from unwanted ones given the language input. Our method enables an end-to-end framework where humans can command the robot to grasp desired objects in a cluttered scene using natural language. Intensive experimental results show the effectiveness of our method in both benchmarking experiments and real-world scenarios, surpassing other baselines. In addition, we demonstrate the practicality of our approach in real-world robotic applications. Our project is available at https://airvlab.github.io/grasp-anything.
title Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.13842