Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Messina, Nicola, Leonardi, Rosario, Ciampi, Luca, Carrara, Fabio, Farinella, Giovanni Maria, Falchi, Fabrizio, Furnari, Antonino
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915647696404480
author Messina, Nicola
Leonardi, Rosario
Ciampi, Luca
Carrara, Fabio
Farinella, Giovanni Maria
Falchi, Fabrizio
Furnari, Antonino
author_facet Messina, Nicola
Leonardi, Rosario
Ciampi, Luca
Carrara, Fabio
Farinella, Giovanni Maria
Falchi, Fabrizio
Furnari, Antonino
contents Pixel-level recognition of objects manipulated by the user from egocentric images enables key applications spanning assistive technologies, industrial safety, and activity monitoring. However, progress in this area is currently hindered by the scarcity of annotated datasets, as existing approaches rely on costly manual labels. In this paper, we propose to learn human-object interaction detection leveraging narrations $\unicode{x2013}$ natural language descriptions of the actions performed by the camera wearer which contain clues about manipulated objects. We introduce Narration-Supervised in-Hand Object Segmentation (NS-iHOS), a novel task where models have to learn to segment in-hand objects by learning from natural-language narrations in a weakly-supervised regime. Narrations are then not employed at inference time. We showcase the potential of the task by proposing Weakly-Supervised In-hand Object Segmentation from Human Narrations (WISH), an end-to-end model distilling knowledge from narrations to learn plausible hand-object associations and enable in-hand object segmentation without using narrations at test time. We benchmark WISH against different baselines based on open-vocabulary object detectors and vision-language models. Experiments on EPIC-Kitchens and Ego4D show that WISH surpasses all baselines, recovering more than 50% of the performance of fully supervised methods, without employing fine-grained pixel-wise annotations. Code and data can be found at https://fpv-iplab.github.io/WISH.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
Messina, Nicola
Leonardi, Rosario
Ciampi, Luca
Carrara, Fabio
Farinella, Giovanni Maria
Falchi, Fabrizio
Furnari, Antonino
Computer Vision and Pattern Recognition
Artificial Intelligence
Pixel-level recognition of objects manipulated by the user from egocentric images enables key applications spanning assistive technologies, industrial safety, and activity monitoring. However, progress in this area is currently hindered by the scarcity of annotated datasets, as existing approaches rely on costly manual labels. In this paper, we propose to learn human-object interaction detection leveraging narrations $\unicode{x2013}$ natural language descriptions of the actions performed by the camera wearer which contain clues about manipulated objects. We introduce Narration-Supervised in-Hand Object Segmentation (NS-iHOS), a novel task where models have to learn to segment in-hand objects by learning from natural-language narrations in a weakly-supervised regime. Narrations are then not employed at inference time. We showcase the potential of the task by proposing Weakly-Supervised In-hand Object Segmentation from Human Narrations (WISH), an end-to-end model distilling knowledge from narrations to learn plausible hand-object associations and enable in-hand object segmentation without using narrations at test time. We benchmark WISH against different baselines based on open-vocabulary object detectors and vision-language models. Experiments on EPIC-Kitchens and Ego4D show that WISH surpasses all baselines, recovering more than 50% of the performance of fully supervised methods, without employing fine-grained pixel-wise annotations. Code and data can be found at https://fpv-iplab.github.io/WISH.
title Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.26004