EgoAnimate: Generating Human Animations from Egocentric top-down Views

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Türkoglu, G. Kutay, Tanke, Julian, Belgacem, Iheb, Markhasin, Lev
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916840505081856
author Türkoglu, G. Kutay
Tanke, Julian
Belgacem, Iheb
Markhasin, Lev
author_facet Türkoglu, G. Kutay
Tanke, Julian
Belgacem, Iheb
Markhasin, Lev
contents An ideal digital telepresence experience requires accurate replication of a person's body, clothing, and movements. To capture and transfer these movements into virtual reality, the egocentric (first-person) perspective can be adopted, which enables the use of a portable and cost-effective device without front-view cameras. However, this viewpoint introduces challenges such as occlusions and distorted body proportions. There are few works reconstructing human appearance from egocentric views, and none use a generative prior-based approach. Some methods create avatars from a single egocentric image during inference, but still rely on multi-view datasets during training. To our knowledge, this is the first study using a generative backbone to reconstruct animatable avatars from egocentric inputs. Based on Stable Diffusion, our method reduces training burden and improves generalizability. Inspired by methods such as SiTH and MagicMan, which perform 360-degree reconstruction from a frontal image, we introduce a pipeline that generates realistic frontal views from occluded top-down images using ControlNet and a Stable Diffusion backbone. Our goal is to convert a single top-down egocentric image into a realistic frontal representation and feed it into an image-to-motion model. This enables generation of avatar motions from minimal input, paving the way for more accessible and generalizable telepresence systems.
format Preprint
id arxiv_https___arxiv_org_abs_2507_09230
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EgoAnimate: Generating Human Animations from Egocentric top-down Views
Türkoglu, G. Kutay
Tanke, Julian
Belgacem, Iheb
Markhasin, Lev
Computer Vision and Pattern Recognition
An ideal digital telepresence experience requires accurate replication of a person's body, clothing, and movements. To capture and transfer these movements into virtual reality, the egocentric (first-person) perspective can be adopted, which enables the use of a portable and cost-effective device without front-view cameras. However, this viewpoint introduces challenges such as occlusions and distorted body proportions. There are few works reconstructing human appearance from egocentric views, and none use a generative prior-based approach. Some methods create avatars from a single egocentric image during inference, but still rely on multi-view datasets during training. To our knowledge, this is the first study using a generative backbone to reconstruct animatable avatars from egocentric inputs. Based on Stable Diffusion, our method reduces training burden and improves generalizability. Inspired by methods such as SiTH and MagicMan, which perform 360-degree reconstruction from a frontal image, we introduce a pipeline that generates realistic frontal views from occluded top-down images using ControlNet and a Stable Diffusion backbone. Our goal is to convert a single top-down egocentric image into a realistic frontal representation and feed it into an image-to-motion model. This enables generation of avatar motions from minimal input, paving the way for more accessible and generalizable telepresence systems.
title EgoAnimate: Generating Human Animations from Egocentric top-down Views
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.09230