EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yuhang, Zhai, Wei, Wang, Chengfeng, Yu, Chengjun, Cao, Yang, Zha, Zheng-Jun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917801626697728
author Yang, Yuhang
Zhai, Wei
Wang, Chengfeng
Yu, Chengjun
Cao, Yang
Zha, Zheng-Jun
author_facet Yang, Yuhang
Zhai, Wei
Wang, Chengfeng
Yu, Chengjun
Cao, Yang
Zha, Zheng-Jun
contents Understanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semantics e.g., ''what'' interaction is occurring, capturing ''where'' the interaction specifically manifests in 3D space is also crucial, which links the perception and operation. Existing methods primarily leverage observations of HOI to capture interaction regions from an exocentric view. However, incomplete observations of interacting parties in the egocentric view introduce ambiguity between visual observations and interaction contents, impairing their efficacy. From the egocentric view, humans integrate the visual cortex, cerebellum, and brain to internalize their intentions and interaction concepts of objects, allowing for the pre-formulation of interactions and making behaviors even when interaction regions are out of sight. In light of this, we propose harmonizing the visual appearance, head motion, and 3D object to excavate the object interaction concept and subject intention, jointly inferring 3D human contact and object affordance from egocentric videos. To achieve this, we present EgoChoir, which links object structures with interaction contexts inherent in appearance and head motion to reveal object affordance, further utilizing it to model human contact. Additionally, a gradient modulation is employed to adopt appropriate clues for capturing interaction regions across various egocentric scenarios. Moreover, 3D contact and affordance are annotated for egocentric videos collected from Ego-Exo4D and GIMO to support the task. Extensive experiments on them demonstrate the effectiveness and superiority of EgoChoir.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13659
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views
Yang, Yuhang
Zhai, Wei
Wang, Chengfeng
Yu, Chengjun
Cao, Yang
Zha, Zheng-Jun
Computer Vision and Pattern Recognition
Understanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semantics e.g., ''what'' interaction is occurring, capturing ''where'' the interaction specifically manifests in 3D space is also crucial, which links the perception and operation. Existing methods primarily leverage observations of HOI to capture interaction regions from an exocentric view. However, incomplete observations of interacting parties in the egocentric view introduce ambiguity between visual observations and interaction contents, impairing their efficacy. From the egocentric view, humans integrate the visual cortex, cerebellum, and brain to internalize their intentions and interaction concepts of objects, allowing for the pre-formulation of interactions and making behaviors even when interaction regions are out of sight. In light of this, we propose harmonizing the visual appearance, head motion, and 3D object to excavate the object interaction concept and subject intention, jointly inferring 3D human contact and object affordance from egocentric videos. To achieve this, we present EgoChoir, which links object structures with interaction contexts inherent in appearance and head motion to reveal object affordance, further utilizing it to model human contact. Additionally, a gradient modulation is employed to adopt appropriate clues for capturing interaction regions across various egocentric scenarios. Moreover, 3D contact and affordance are annotated for egocentric videos collected from Ego-Exo4D and GIMO to support the task. Extensive experiments on them demonstrate the effectiveness and superiority of EgoChoir.
title EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.13659