Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Materia, Daniele, Ragusa, Francesco, Farinella, Giovanni Maria |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
by: Leonardi, Rosario, et al.
Published: (2026)
by: Leonardi, Rosario, et al.
Published: (2026)
Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction Detection
by: Leonardi, Rosario, et al.
Published: (2026)
by: Leonardi, Rosario, et al.
Published: (2026)
StillFast: An End-to-End Approach for Short-Term Object Interaction Anticipation
by: Ragusa, Francesco, et al.
Published: (2023)
by: Ragusa, Francesco, et al.
Published: (2023)
Exploiting Multimodal Synthetic Data for Egocentric Human-Object Interaction Detection in an Industrial Scenario
by: Leonardi, Rosario, et al.
Published: (2023)
by: Leonardi, Rosario, et al.
Published: (2023)
GlovEgo-HOI: Bridging the Synthetic-to-Real Gap for Industrial Egocentric Human-Object Interaction Detection
by: Spoto, Alfio, et al.
Published: (2026)
by: Spoto, Alfio, et al.
Published: (2026)
Are Synthetic Data Useful for Egocentric Hand-Object Interaction Detection?
by: Leonardi, Rosario, et al.
Published: (2023)
by: Leonardi, Rosario, et al.
Published: (2023)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
by: Mazzamuto, Michele, et al.
Published: (2024)
by: Mazzamuto, Michele, et al.
Published: (2024)
SignIT: A Comprehensive Dataset and Multimodal Analysis for Italian Sign Language Recognition
by: Micieli, Alessia, et al.
Published: (2025)
by: Micieli, Alessia, et al.
Published: (2025)
ZARRIO @ Ego4D Short Term Object Interaction Anticipation Challenge: Leveraging Affordances and Attention-based models for STA
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
An Outlook into the Future of Egocentric Vision
by: Plizzari, Chiara, et al.
Published: (2023)
by: Plizzari, Chiara, et al.
Published: (2023)
AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2025)
by: Seminara, Luigi, et al.
Published: (2025)
Integrating Affordances and Attention models for Short-Term Object Interaction Anticipation
by: Labadia, Lorenzo Mur, et al.
Published: (2026)
by: Labadia, Lorenzo Mur, et al.
Published: (2026)
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos
by: Seminara, Luigi, et al.
Published: (2024)
by: Seminara, Luigi, et al.
Published: (2024)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
A Real-Time System for Egocentric Hand-Object Interaction Detection in Industrial Domains
by: Finocchiaro, Antonio, et al.
Published: (2025)
by: Finocchiaro, Antonio, et al.
Published: (2025)
Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
by: Ragusa, Francesco, et al.
Published: (2025)
by: Ragusa, Francesco, et al.
Published: (2025)
Mamba-OTR: a Mamba-based Solution for Online Take and Release Detection from Untrimmed Egocentric Video
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
by: Lai, Bolin, et al.
Published: (2023)
by: Lai, Bolin, et al.
Published: (2023)
Synchronization is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs
by: Quattrocchi, Camillo, et al.
Published: (2023)
by: Quattrocchi, Camillo, et al.
Published: (2023)
Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
by: Messina, Nicola, et al.
Published: (2025)
by: Messina, Nicola, et al.
Published: (2025)
Semantically Guided Action Anticipation
by: Diko, Anxhelo, et al.
Published: (2024)
by: Diko, Anxhelo, et al.
Published: (2024)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
by: Luo, Yuanhao, et al.
Published: (2026)
by: Luo, Yuanhao, et al.
Published: (2026)
ENIGMA-360: An Ego-Exo Dataset for Human Behavior Understanding in Industrial Scenarios
by: Ragusa, Francesco, et al.
Published: (2026)
by: Ragusa, Francesco, et al.
Published: (2026)
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
by: Lall, Vishakha, et al.
Published: (2025)
by: Lall, Vishakha, et al.
Published: (2025)
Interaction Region Visual Transformer for Egocentric Action Anticipation
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs
by: Rodin, Ivan, et al.
Published: (2025)
by: Rodin, Ivan, et al.
Published: (2025)
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
by: Manigrasso, Zaira, et al.
Published: (2024)
by: Manigrasso, Zaira, et al.
Published: (2024)
Calisthenics Skills Temporal Video Segmentation
by: Finocchiaro, Antonio, et al.
Published: (2025)
by: Finocchiaro, Antonio, et al.
Published: (2025)
ProSkill: Segment-Level Skill Assessment in Procedural Videos
by: Mazzamuto, Michele, et al.
Published: (2026)
by: Mazzamuto, Michele, et al.
Published: (2026)
Ego-METAS: Egocentric online Multimodal Energy-efficient Temporal Action Segmentation benchmark
by: Santos-Villafranca, Maria, et al.
Published: (2026)
by: Santos-Villafranca, Maria, et al.
Published: (2026)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
Egocentric Human-Object Interaction Detection: A New Benchmark and Method
by: Deng, Kunyuan, et al.
Published: (2025)
by: Deng, Kunyuan, et al.
Published: (2025)
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
by: Peng, Taiying, et al.
Published: (2025)
by: Peng, Taiying, et al.
Published: (2025)
Personalized Federated Learning for Egocentric Video Gaze Estimation with Comprehensive Parameter Frezzing
by: Feng, Yuhu, et al.
Published: (2025)
by: Feng, Yuhu, et al.
Published: (2025)
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024)
by: Manousaki, Victoria, et al.
Published: (2024)
PEAR: Phrase-Based Hand-Object Interaction Anticipation
by: Zhang, Zichen, et al.
Published: (2024)
by: Zhang, Zichen, et al.
Published: (2024)
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
by: Xu, Liang, et al.
Published: (2025)
by: Xu, Liang, et al.
Published: (2025)
HOIMotion: Forecasting Human Motion During Human-Object Interactions Using Egocentric 3D Object Bounding Boxes
by: Hu, Zhiming, et al.
Published: (2024)
by: Hu, Zhiming, et al.
Published: (2024)
Similar Items
-
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
by: Leonardi, Rosario, et al.
Published: (2026) -
Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction Detection
by: Leonardi, Rosario, et al.
Published: (2026) -
StillFast: An End-to-End Approach for Short-Term Object Interaction Anticipation
by: Ragusa, Francesco, et al.
Published: (2023) -
Exploiting Multimodal Synthetic Data for Egocentric Human-Object Interaction Detection in an Industrial Scenario
by: Leonardi, Rosario, et al.
Published: (2023) -
GlovEgo-HOI: Bridging the Synthetic-to-Real Gap for Industrial Egocentric Human-Object Interaction Detection
by: Spoto, Alfio, et al.
Published: (2026)