Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
Fuente:
arXiv
Saved in:
| Main Authors: | Messina, Nicola, Leonardi, Rosario, Ciampi, Luca, Carrara, Fabio, Farinella, Giovanni Maria, Falchi, Fabrizio, Furnari, Antonino |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are Synthetic Data Useful for Egocentric Hand-Object Interaction Detection?
by: Leonardi, Rosario, et al.
Published: (2023)
by: Leonardi, Rosario, et al.
Published: (2023)
Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction Detection
by: Leonardi, Rosario, et al.
Published: (2026)
by: Leonardi, Rosario, et al.
Published: (2026)
Exploiting Multimodal Synthetic Data for Egocentric Human-Object Interaction Detection in an Industrial Scenario
by: Leonardi, Rosario, et al.
Published: (2023)
by: Leonardi, Rosario, et al.
Published: (2023)
A Real-Time System for Egocentric Hand-Object Interaction Detection in Industrial Domains
by: Finocchiaro, Antonio, et al.
Published: (2025)
by: Finocchiaro, Antonio, et al.
Published: (2025)
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2024)
by: Seminara, Luigi, et al.
Published: (2024)
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2025)
by: Seminara, Luigi, et al.
Published: (2025)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
by: Mazzamuto, Michele, et al.
Published: (2024)
by: Mazzamuto, Michele, et al.
Published: (2024)
Is CLIP the main roadblock for fine-grained open-world perception?
by: Bianchi, Lorenzo, et al.
Published: (2024)
by: Bianchi, Lorenzo, et al.
Published: (2024)
Mamba-OTR: a Mamba-based Solution for Online Take and Release Detection from Untrimmed Egocentric Video
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)
Calisthenics Skills Temporal Video Segmentation
by: Finocchiaro, Antonio, et al.
Published: (2025)
by: Finocchiaro, Antonio, et al.
Published: (2025)
StillFast: An End-to-End Approach for Short-Term Object Interaction Anticipation
by: Ragusa, Francesco, et al.
Published: (2023)
by: Ragusa, Francesco, et al.
Published: (2023)
Semi-Supervised Biomedical Image Segmentation via Diffusion Models and Teacher-Student Co-Training
by: Ciampi, Luca, et al.
Published: (2025)
by: Ciampi, Luca, et al.
Published: (2025)
GlovEgo-HOI: Bridging the Synthetic-to-Real Gap for Industrial Egocentric Human-Object Interaction Detection
by: Spoto, Alfio, et al.
Published: (2026)
by: Spoto, Alfio, et al.
Published: (2026)
Efficient Calisthenics Skills Classification through Foreground Instance Selection and Depth Estimation
by: Finocchiaro, Antonio, et al.
Published: (2025)
by: Finocchiaro, Antonio, et al.
Published: (2025)
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
by: Barsellotti, Luca, et al.
Published: (2024)
by: Barsellotti, Luca, et al.
Published: (2024)
Does it Really Count? Assessing Semantic Grounding in Text-Guided Class-Agnostic Counting
by: Pacini, Giacomo, et al.
Published: (2026)
by: Pacini, Giacomo, et al.
Published: (2026)
Ego-METAS: Egocentric online Multimodal Energy-efficient Temporal Action Segmentation benchmark
by: Santos-Villafranca, Maria, et al.
Published: (2026)
by: Santos-Villafranca, Maria, et al.
Published: (2026)
Synchronization is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs
by: Quattrocchi, Camillo, et al.
Published: (2023)
by: Quattrocchi, Camillo, et al.
Published: (2023)
The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understanding
by: Bianchi, Lorenzo, et al.
Published: (2023)
by: Bianchi, Lorenzo, et al.
Published: (2023)
Mind the Prompt: A Novel Benchmark for Prompt-based Class-Agnostic Counting
by: Ciampi, Luca, et al.
Published: (2024)
by: Ciampi, Luca, et al.
Published: (2024)
CountingDINO: A Training-free Pipeline for Class-Agnostic Counting using Unsupervised Backbones
by: Pacini, Giacomo, et al.
Published: (2025)
by: Pacini, Giacomo, et al.
Published: (2025)
Biologically-inspired Semi-supervised Semantic Segmentation for Biomedical Imaging
by: Ciampi, Luca, et al.
Published: (2024)
by: Ciampi, Luca, et al.
Published: (2024)
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
by: Manigrasso, Zaira, et al.
Published: (2024)
by: Manigrasso, Zaira, et al.
Published: (2024)
How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?
by: Lando, Giuseppe, et al.
Published: (2025)
by: Lando, Giuseppe, et al.
Published: (2025)
One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
by: Bianchi, Lorenzo, et al.
Published: (2025)
by: Bianchi, Lorenzo, et al.
Published: (2025)
An Outlook into the Future of Egocentric Vision
by: Plizzari, Chiara, et al.
Published: (2023)
by: Plizzari, Chiara, et al.
Published: (2023)
Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
by: Ragusa, Francesco, et al.
Published: (2025)
by: Ragusa, Francesco, et al.
Published: (2025)
EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision
by: Forte, Rosario, et al.
Published: (2026)
by: Forte, Rosario, et al.
Published: (2026)
ProSkill: Segment-Level Skill Assessment in Procedural Videos
by: Mazzamuto, Michele, et al.
Published: (2026)
by: Mazzamuto, Michele, et al.
Published: (2026)
Semantically Guided Action Anticipation
by: Diko, Anxhelo, et al.
Published: (2024)
by: Diko, Anxhelo, et al.
Published: (2024)
EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs
by: Rodin, Ivan, et al.
Published: (2025)
by: Rodin, Ivan, et al.
Published: (2025)
AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos
by: Materia, Daniele, et al.
Published: (2026)
by: Materia, Daniele, et al.
Published: (2026)
Joint-Dataset Learning and Cross-Consistent Regularization for Text-to-Motion Retrieval
by: Messina, Nicola, et al.
Published: (2024)
by: Messina, Nicola, et al.
Published: (2024)
ENIGMA-360: An Ego-Exo Dataset for Human Behavior Understanding in Industrial Scenarios
by: Ragusa, Francesco, et al.
Published: (2026)
by: Ragusa, Francesco, et al.
Published: (2026)
Integrating Affordances and Attention models for Short-Term Object Interaction Anticipation
by: Labadia, Lorenzo Mur, et al.
Published: (2026)
by: Labadia, Lorenzo Mur, et al.
Published: (2026)
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
by: Leonardi, Rosario, et al.
Published: (2026)
by: Leonardi, Rosario, et al.
Published: (2026)
Maybe you are looking for CroQS: Cross-modal Query Suggestion for Text-to-Image Retrieval
by: Pacini, Giacomo, et al.
Published: (2024)
by: Pacini, Giacomo, et al.
Published: (2024)
Exploring Multimodal LMMs for Online Episodic Memory Question Answering on the Edge
by: Lando, Giuseppe, et al.
Published: (2026)
by: Lando, Giuseppe, et al.
Published: (2026)
TI-PREGO: Chain of Thought and In-Context Learning for Online Mistake Detection in PRocedural EGOcentric Videos
by: Plini, Leonardo, et al.
Published: (2024)
by: Plini, Leonardo, et al.
Published: (2024)
Similar Items
-
Are Synthetic Data Useful for Egocentric Hand-Object Interaction Detection?
by: Leonardi, Rosario, et al.
Published: (2023) -
Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction Detection
by: Leonardi, Rosario, et al.
Published: (2026) -
Exploiting Multimodal Synthetic Data for Egocentric Human-Object Interaction Detection in an Industrial Scenario
by: Leonardi, Rosario, et al.
Published: (2023) -
A Real-Time System for Egocentric Hand-Object Interaction Detection in Industrial Domains
by: Finocchiaro, Antonio, et al.
Published: (2025) -
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2024)