Object segmentation in the wild with foundation models: application to vision assisted neuro-prostheses for upper limbs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Atoki, Bolutife, Benois-Pineau, Jenny, Péteri, Renaud, Baldacci, Fabien, de Rugy, Aymar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912499893272576
author Atoki, Bolutife
Benois-Pineau, Jenny
Péteri, Renaud
Baldacci, Fabien
de Rugy, Aymar
author_facet Atoki, Bolutife
Benois-Pineau, Jenny
Péteri, Renaud
Baldacci, Fabien
de Rugy, Aymar
contents In this work, we address the problem of semantic object segmentation using foundation models. We investigate whether foundation models, trained on a large number and variety of objects, can perform object segmentation without fine-tuning on specific images containing everyday objects, but in highly cluttered visual scenes. The ''in the wild'' context is driven by the target application of vision guided upper limb neuroprostheses. We propose a method for generating prompts based on gaze fixations to guide the Segment Anything Model (SAM) in our segmentation scenario, and fine-tune it on egocentric visual data. Evaluation results of our approach show an improvement of the IoU segmentation quality metric by up to 0.51 points on real-world challenging data of Grasping-in-the-Wild corpus which is made available on the RoboFlow Platform (https://universe.roboflow.com/iwrist/grasping-in-the-wild)
format Preprint
id arxiv_https___arxiv_org_abs_2507_18517
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Object segmentation in the wild with foundation models: application to vision assisted neuro-prostheses for upper limbs
Atoki, Bolutife
Benois-Pineau, Jenny
Péteri, Renaud
Baldacci, Fabien
de Rugy, Aymar
Computer Vision and Pattern Recognition
In this work, we address the problem of semantic object segmentation using foundation models. We investigate whether foundation models, trained on a large number and variety of objects, can perform object segmentation without fine-tuning on specific images containing everyday objects, but in highly cluttered visual scenes. The ''in the wild'' context is driven by the target application of vision guided upper limb neuroprostheses. We propose a method for generating prompts based on gaze fixations to guide the Segment Anything Model (SAM) in our segmentation scenario, and fine-tune it on egocentric visual data. Evaluation results of our approach show an improvement of the IoU segmentation quality metric by up to 0.51 points on real-world challenging data of Grasping-in-the-Wild corpus which is made available on the RoboFlow Platform (https://universe.roboflow.com/iwrist/grasping-in-the-wild)
title Object segmentation in the wild with foundation models: application to vision assisted neuro-prostheses for upper limbs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.18517