Learning Efficient and Generalizable Human Representation with Human Gaussian Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Yifan, Zhang, Shengjun, Dai, Chensheng, Chen, Yang, Liu, Hao, Li, Chen, Duan, Yueqi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916862280859648
author Liu, Yifan
Zhang, Shengjun
Dai, Chensheng
Chen, Yang
Liu, Hao
Li, Chen
Duan, Yueqi
author_facet Liu, Yifan
Zhang, Shengjun
Dai, Chensheng
Chen, Yang
Liu, Hao
Li, Chen
Duan, Yueqi
contents Modeling animatable human avatars from videos is a long-standing and challenging problem. While conventional methods require per-instance optimization, recent feed-forward methods have been proposed to generate 3D Gaussians with a learnable network. However, these methods predict Gaussians for each frame independently, without fully capturing the relations of Gaussians from different timestamps. To address this, we propose Human Gaussian Graph to model the connection between predicted Gaussians and human SMPL mesh, so that we can leverage information from all frames to recover an animatable human representation. Specifically, the Human Gaussian Graph contains dual layers where Gaussians are the first layer nodes and mesh vertices serve as the second layer nodes. Based on this structure, we further propose the intra-node operation to aggregate various Gaussians connected to one mesh vertex, and inter-node operation to support message passing among mesh node neighbors. Experimental results on novel view synthesis and novel pose animation demonstrate the efficiency and generalization of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18758
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Efficient and Generalizable Human Representation with Human Gaussian Model
Liu, Yifan
Zhang, Shengjun
Dai, Chensheng
Chen, Yang
Liu, Hao
Li, Chen
Duan, Yueqi
Computer Vision and Pattern Recognition
Modeling animatable human avatars from videos is a long-standing and challenging problem. While conventional methods require per-instance optimization, recent feed-forward methods have been proposed to generate 3D Gaussians with a learnable network. However, these methods predict Gaussians for each frame independently, without fully capturing the relations of Gaussians from different timestamps. To address this, we propose Human Gaussian Graph to model the connection between predicted Gaussians and human SMPL mesh, so that we can leverage information from all frames to recover an animatable human representation. Specifically, the Human Gaussian Graph contains dual layers where Gaussians are the first layer nodes and mesh vertices serve as the second layer nodes. Based on this structure, we further propose the intra-node operation to aggregate various Gaussians connected to one mesh vertex, and inter-node operation to support message passing among mesh node neighbors. Experimental results on novel view synthesis and novel pose animation demonstrate the efficiency and generalization of our method.
title Learning Efficient and Generalizable Human Representation with Human Gaussian Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.18758