Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Haosheng, Mao, Weixin, Deng, Weipeng, Meng, Chenyu, Fan, Haoqiang, Wang, Tiancai, Osamu, Yoshie, Tan, Ping, Wang, Hongan, Deng, Xiaoming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910993743872000
author Li, Haosheng
Mao, Weixin
Deng, Weipeng
Meng, Chenyu
Fan, Haoqiang
Wang, Tiancai
Osamu, Yoshie
Tan, Ping
Wang, Hongan
Deng, Xiaoming
author_facet Li, Haosheng
Mao, Weixin
Deng, Weipeng
Meng, Chenyu
Fan, Haoqiang
Wang, Tiancai
Osamu, Yoshie
Tan, Ping
Wang, Hongan
Deng, Xiaoming
contents Multi-hand semantic grasp generation aims to generate feasible and semantically appropriate grasp poses for different robotic hands based on natural language instructions. Although the task is highly valuable, due to the lack of multihand grasp datasets with fine-grained contact description between robotic hands and objects, it is still a long-standing difficult task. In this paper, we present Multi-GraspSet, the first large-scale multi-hand grasp dataset with automatically contact annotations. Based on Multi-GraspSet, we propose Multi-GraspLLM, a unified language-guided grasp generation framework, which leverages large language models (LLM) to handle variable-length sequences, generating grasp poses for diverse robotic hands in a single unified architecture. Multi-GraspLLM first aligns the encoded point cloud features and text features into a unified semantic space. It then generates grasp bin tokens that are subsequently converted into grasp pose for each robotic hand via hand-aware linear mapping. The experimental results demonstrate that our approach significantly outperforms existing methods in both real-world experiments and simulator. More information can be found on our project page https://multi-graspllm.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08468
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation
Li, Haosheng
Mao, Weixin
Deng, Weipeng
Meng, Chenyu
Fan, Haoqiang
Wang, Tiancai
Osamu, Yoshie
Tan, Ping
Wang, Hongan
Deng, Xiaoming
Robotics
Computer Vision and Pattern Recognition
Multi-hand semantic grasp generation aims to generate feasible and semantically appropriate grasp poses for different robotic hands based on natural language instructions. Although the task is highly valuable, due to the lack of multihand grasp datasets with fine-grained contact description between robotic hands and objects, it is still a long-standing difficult task. In this paper, we present Multi-GraspSet, the first large-scale multi-hand grasp dataset with automatically contact annotations. Based on Multi-GraspSet, we propose Multi-GraspLLM, a unified language-guided grasp generation framework, which leverages large language models (LLM) to handle variable-length sequences, generating grasp poses for diverse robotic hands in a single unified architecture. Multi-GraspLLM first aligns the encoded point cloud features and text features into a unified semantic space. It then generates grasp bin tokens that are subsequently converted into grasp pose for each robotic hand via hand-aware linear mapping. The experimental results demonstrate that our approach significantly outperforms existing methods in both real-world experiments and simulator. More information can be found on our project page https://multi-graspllm.github.io.
title Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.08468