Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Muttaqien, Muhammad A., Motoda, Tomohiro, Hanai, Ryo, Yukiyasu, Domae
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916634875133952
author Muttaqien, Muhammad A.
Motoda, Tomohiro
Hanai, Ryo
Yukiyasu, Domae
author_facet Muttaqien, Muhammad A.
Motoda, Tomohiro
Hanai, Ryo
Yukiyasu, Domae
contents This paper introduces a novel pipeline to enhance the precision of object masking for robotic manipulation within the specific domain of masking products in convenience stores. The approach integrates two advanced AI models, CLIP and SAM, focusing on their synergistic combination and the effective use of multimodal data (image and text). Emphasis is placed on utilizing gradient-based attention mechanisms and customized datasets to fine-tune performance. While CLIP, SAM, and Grad- CAM are established components, their integration within this structured pipeline represents a significant contribution to the field. The resulting segmented masks, generated through this combined approach, can be effectively utilized as inputs for robotic systems, enabling more precise and adaptive object manipulation in the context of convenience store products.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation
Muttaqien, Muhammad A.
Motoda, Tomohiro
Hanai, Ryo
Yukiyasu, Domae
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
This paper introduces a novel pipeline to enhance the precision of object masking for robotic manipulation within the specific domain of masking products in convenience stores. The approach integrates two advanced AI models, CLIP and SAM, focusing on their synergistic combination and the effective use of multimodal data (image and text). Emphasis is placed on utilizing gradient-based attention mechanisms and customized datasets to fine-tune performance. While CLIP, SAM, and Grad- CAM are established components, their integration within this structured pipeline represents a significant contribution to the field. The resulting segmented masks, generated through this combined approach, can be effectively utilized as inputs for robotic systems, enabling more precise and adaptive object manipulation in the context of convenience store products.
title Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.18842