HRM^2Avatar: High-Fidelity Real-Time Mobile Avatars from Monocular Phone Scans

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Chao, Jia, Shenghao, Liu, Jinhui, Zhang, Yong, Zhu, Liangchao, Yang, Zhonglei, Ma, Jinze, Niu, Chaoyue, Lv, Chengfei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912675074670592
author Shi, Chao
Jia, Shenghao
Liu, Jinhui
Zhang, Yong
Zhu, Liangchao
Yang, Zhonglei
Ma, Jinze
Niu, Chaoyue
Lv, Chengfei
author_facet Shi, Chao
Jia, Shenghao
Liu, Jinhui
Zhang, Yong
Zhu, Liangchao
Yang, Zhonglei
Ma, Jinze
Niu, Chaoyue
Lv, Chengfei
contents We present HRM$^2$Avatar, a framework for creating high-fidelity avatars from monocular phone scans, which can be rendered and animated in real time on mobile devices. Monocular capture with smartphones provides a low-cost alternative to studio-grade multi-camera rigs, making avatar digitization accessible to non-expert users. Reconstructing high-fidelity avatars from single-view video sequences poses challenges due to limited visual and geometric data. To address these limitations, at the data level, our method leverages two types of data captured with smartphones: static pose sequences for texture reconstruction and dynamic motion sequences for learning pose-dependent deformations and lighting changes. At the representation level, we employ a lightweight yet expressive representation to reconstruct high-fidelity digital humans from sparse monocular data. We extract garment meshes from monocular data to model clothing deformations effectively, and attach illumination-aware Gaussians to the mesh surface, enabling high-fidelity rendering and capturing pose-dependent lighting. This representation efficiently learns high-resolution and dynamic information from monocular data, enabling the creation of detailed avatars. At the rendering level, real-time performance is critical for animating high-fidelity avatars in AR/VR, social gaming, and on-device creation. Our GPU-driven rendering pipeline delivers 120 FPS on mobile devices and 90 FPS on standalone VR devices at 2K resolution, over $2.7\times$ faster than representative mobile-engine baselines. Experiments show that HRM$^2$Avatar delivers superior visual realism and real-time interactivity, outperforming state-of-the-art monocular methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13587
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HRM^2Avatar: High-Fidelity Real-Time Mobile Avatars from Monocular Phone Scans
Shi, Chao
Jia, Shenghao
Liu, Jinhui
Zhang, Yong
Zhu, Liangchao
Yang, Zhonglei
Ma, Jinze
Niu, Chaoyue
Lv, Chengfei
Graphics
We present HRM$^2$Avatar, a framework for creating high-fidelity avatars from monocular phone scans, which can be rendered and animated in real time on mobile devices. Monocular capture with smartphones provides a low-cost alternative to studio-grade multi-camera rigs, making avatar digitization accessible to non-expert users. Reconstructing high-fidelity avatars from single-view video sequences poses challenges due to limited visual and geometric data. To address these limitations, at the data level, our method leverages two types of data captured with smartphones: static pose sequences for texture reconstruction and dynamic motion sequences for learning pose-dependent deformations and lighting changes. At the representation level, we employ a lightweight yet expressive representation to reconstruct high-fidelity digital humans from sparse monocular data. We extract garment meshes from monocular data to model clothing deformations effectively, and attach illumination-aware Gaussians to the mesh surface, enabling high-fidelity rendering and capturing pose-dependent lighting. This representation efficiently learns high-resolution and dynamic information from monocular data, enabling the creation of detailed avatars. At the rendering level, real-time performance is critical for animating high-fidelity avatars in AR/VR, social gaming, and on-device creation. Our GPU-driven rendering pipeline delivers 120 FPS on mobile devices and 90 FPS on standalone VR devices at 2K resolution, over $2.7\times$ faster than representative mobile-engine baselines. Experiments show that HRM$^2$Avatar delivers superior visual realism and real-time interactivity, outperforming state-of-the-art monocular methods.
title HRM^2Avatar: High-Fidelity Real-Time Mobile Avatars from Monocular Phone Scans
topic Graphics
url https://arxiv.org/abs/2510.13587