GraspSAM: When Segment Anything Model Meets Grasp Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Noh, Sangjun, Kim, Jongwon, Nam, Dongwoo, Back, Seunghyeok, Kang, Raeyoung, Lee, Kyoobin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916405428879360
author Noh, Sangjun
Kim, Jongwon
Nam, Dongwoo
Back, Seunghyeok
Kang, Raeyoung
Lee, Kyoobin
author_facet Noh, Sangjun
Kim, Jongwon
Nam, Dongwoo
Back, Seunghyeok
Kang, Raeyoung
Lee, Kyoobin
contents Grasp detection requires flexibility to handle objects of various shapes without relying on prior knowledge of the object, while also offering intuitive, user-guided control. This paper introduces GraspSAM, an innovative extension of the Segment Anything Model (SAM), designed for prompt-driven and category-agnostic grasp detection. Unlike previous methods, which are often limited by small-scale training data, GraspSAM leverages the large-scale training and prompt-based segmentation capabilities of SAM to efficiently support both target-object and category-agnostic grasping. By utilizing adapters, learnable token embeddings, and a lightweight modified decoder, GraspSAM requires minimal fine-tuning to integrate object segmentation and grasp prediction into a unified framework. The model achieves state-of-the-art (SOTA) performance across multiple datasets, including Jacquard, Grasp-Anything, and Grasp-Anything++. Extensive experiments demonstrate the flexibility of GraspSAM in handling different types of prompts (such as points, boxes, and language), highlighting its robustness and effectiveness in real-world robotic applications.
format Preprint
id arxiv_https___arxiv_org_abs_2409_12521
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GraspSAM: When Segment Anything Model Meets Grasp Detection
Noh, Sangjun
Kim, Jongwon
Nam, Dongwoo
Back, Seunghyeok
Kang, Raeyoung
Lee, Kyoobin
Robotics
Systems and Control
Grasp detection requires flexibility to handle objects of various shapes without relying on prior knowledge of the object, while also offering intuitive, user-guided control. This paper introduces GraspSAM, an innovative extension of the Segment Anything Model (SAM), designed for prompt-driven and category-agnostic grasp detection. Unlike previous methods, which are often limited by small-scale training data, GraspSAM leverages the large-scale training and prompt-based segmentation capabilities of SAM to efficiently support both target-object and category-agnostic grasping. By utilizing adapters, learnable token embeddings, and a lightweight modified decoder, GraspSAM requires minimal fine-tuning to integrate object segmentation and grasp prediction into a unified framework. The model achieves state-of-the-art (SOTA) performance across multiple datasets, including Jacquard, Grasp-Anything, and Grasp-Anything++. Extensive experiments demonstrate the flexibility of GraspSAM in handling different types of prompts (such as points, boxes, and language), highlighting its robustness and effectiveness in real-world robotic applications.
title GraspSAM: When Segment Anything Model Meets Grasp Detection
topic Robotics
Systems and Control
url https://arxiv.org/abs/2409.12521