FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Camiletto, Andrea Boscolo, Wang, Jian, Alvarado, Eduardo, Dabral, Rishabh, Beeler, Thabo, Habermann, Marc, Theobalt, Christian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909557674999808
author Camiletto, Andrea Boscolo
Wang, Jian
Alvarado, Eduardo
Dabral, Rishabh
Beeler, Thabo
Habermann, Marc
Theobalt, Christian
author_facet Camiletto, Andrea Boscolo
Wang, Jian
Alvarado, Eduardo
Dabral, Rishabh
Beeler, Thabo
Habermann, Marc
Theobalt, Christian
contents Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on synthetic pretraining and struggle to generate smooth and accurate predictions in real-world settings, particularly for lower limbs. Our work addresses these limitations by introducing a lightweight VR-based data collection setup with on-board, real-time 6D pose tracking. Using this setup, we collected the most extensive real-world dataset for ego-facing ego-mounted cameras to date in size and motion variability. Effectively integrating this multimodal input -- device pose and camera feeds -- is challenging due to the differing characteristics of each data source. To address this, we propose FRAME, a simple yet effective architecture that combines device pose and camera feeds for state-of-the-art body pose prediction through geometrically sound multimodal integration and can run at 300 FPS on modern hardware. Lastly, we showcase a novel training strategy to enhance the model's generalization capabilities. Our approach exploits the problem's geometric properties, yielding high-quality motion capture free from common artifacts in prior works. Qualitative and quantitative evaluations, along with extensive comparisons, demonstrate the effectiveness of our method. Data, code, and CAD designs will be available at https://vcai.mpi-inf.mpg.de/projects/FRAME/
format Preprint
id arxiv_https___arxiv_org_abs_2503_23094
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video
Camiletto, Andrea Boscolo
Wang, Jian
Alvarado, Eduardo
Dabral, Rishabh
Beeler, Thabo
Habermann, Marc
Theobalt, Christian
Computer Vision and Pattern Recognition
Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on synthetic pretraining and struggle to generate smooth and accurate predictions in real-world settings, particularly for lower limbs. Our work addresses these limitations by introducing a lightweight VR-based data collection setup with on-board, real-time 6D pose tracking. Using this setup, we collected the most extensive real-world dataset for ego-facing ego-mounted cameras to date in size and motion variability. Effectively integrating this multimodal input -- device pose and camera feeds -- is challenging due to the differing characteristics of each data source. To address this, we propose FRAME, a simple yet effective architecture that combines device pose and camera feeds for state-of-the-art body pose prediction through geometrically sound multimodal integration and can run at 300 FPS on modern hardware. Lastly, we showcase a novel training strategy to enhance the model's generalization capabilities. Our approach exploits the problem's geometric properties, yielding high-quality motion capture free from common artifacts in prior works. Qualitative and quantitative evaluations, along with extensive comparisons, demonstrate the effectiveness of our method. Data, code, and CAD designs will be available at https://vcai.mpi-inf.mpg.de/projects/FRAME/
title FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.23094