Lightweight Language-driven Grasp Detection using Conditional Consistency Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nguyen, Nghia, Vu, Minh Nhat, Huang, Baoru, Vuong, An, Le, Ngan, Vo, Thieu, Nguyen, Anh
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917733149442048
author Nguyen, Nghia
Vu, Minh Nhat
Huang, Baoru
Vuong, An
Le, Ngan
Vo, Thieu
Nguyen, Anh
author_facet Nguyen, Nghia
Vu, Minh Nhat
Huang, Baoru
Vuong, An
Le, Ngan
Vo, Thieu
Nguyen, Anh
contents Language-driven grasp detection is a fundamental yet challenging task in robotics with various industrial applications. In this work, we present a new approach for language-driven grasp detection that leverages the concept of lightweight diffusion models to achieve fast inference time. By integrating diffusion processes with grasping prompts in natural language, our method can effectively encode visual and textual information, enabling more accurate and versatile grasp positioning that aligns well with the text query. To overcome the long inference time problem in diffusion models, we leverage the image and text features as the condition in the consistency model to reduce the number of denoising timesteps during inference. The intensive experimental results show that our method outperforms other recent grasp detection methods and lightweight diffusion models by a clear margin. We further validate our method in real-world robotic experiments to demonstrate its fast inference time capability.
format Preprint
id arxiv_https___arxiv_org_abs_2407_17967
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Lightweight Language-driven Grasp Detection using Conditional Consistency Model
Nguyen, Nghia
Vu, Minh Nhat
Huang, Baoru
Vuong, An
Le, Ngan
Vo, Thieu
Nguyen, Anh
Robotics
Computer Vision and Pattern Recognition
Language-driven grasp detection is a fundamental yet challenging task in robotics with various industrial applications. In this work, we present a new approach for language-driven grasp detection that leverages the concept of lightweight diffusion models to achieve fast inference time. By integrating diffusion processes with grasping prompts in natural language, our method can effectively encode visual and textual information, enabling more accurate and versatile grasp positioning that aligns well with the text query. To overcome the long inference time problem in diffusion models, we leverage the image and text features as the condition in the consistency model to reduce the number of denoising timesteps during inference. The intensive experimental results show that our method outperforms other recent grasp detection methods and lightweight diffusion models by a clear margin. We further validate our method in real-world robotic experiments to demonstrate its fast inference time capability.
title Lightweight Language-driven Grasp Detection using Conditional Consistency Model
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.17967