Dynamic Avatar-Scene Rendering from Human-centric Context

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Wenqing, Yang, Haosen, Kittler, Josef, Zhu, Xiatian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911263844466688
author Wang, Wenqing
Yang, Haosen
Kittler, Josef
Zhu, Xiatian
author_facet Wang, Wenqing
Yang, Haosen
Kittler, Josef
Zhu, Xiatian
contents Reconstructing dynamic humans interacting with real-world environments from monocular videos is an important and challenging task. Despite considerable progress in 4D neural rendering, existing approaches either model dynamic scenes holistically or model scenes and backgrounds separately aim to introduce parametric human priors. However, these approaches either neglect distinct motion characteristics of various components in scene especially human, leading to incomplete reconstructions, or ignore the information exchange between the separately modeled components, resulting in spatial inconsistencies and visual artifacts at human-scene boundaries. To address this, we propose {\bf Separate-then-Map} (StM) strategy that introduces a dedicated information mapping mechanism to bridge separately defined and optimized models. Our method employs a shared transformation function for each Gaussian attribute to unify separately modeled components, enhancing computational efficiency by avoiding exhaustive pairwise interactions while ensuring spatial and visual coherence between humans and their surroundings. Extensive experiments on monocular video datasets demonstrate that StM significantly outperforms existing state-of-the-art methods in both visual quality and rendering accuracy, particularly at challenging human-scene interaction boundaries.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10539
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Avatar-Scene Rendering from Human-centric Context
Wang, Wenqing
Yang, Haosen
Kittler, Josef
Zhu, Xiatian
Computer Vision and Pattern Recognition
Reconstructing dynamic humans interacting with real-world environments from monocular videos is an important and challenging task. Despite considerable progress in 4D neural rendering, existing approaches either model dynamic scenes holistically or model scenes and backgrounds separately aim to introduce parametric human priors. However, these approaches either neglect distinct motion characteristics of various components in scene especially human, leading to incomplete reconstructions, or ignore the information exchange between the separately modeled components, resulting in spatial inconsistencies and visual artifacts at human-scene boundaries. To address this, we propose {\bf Separate-then-Map} (StM) strategy that introduces a dedicated information mapping mechanism to bridge separately defined and optimized models. Our method employs a shared transformation function for each Gaussian attribute to unify separately modeled components, enhancing computational efficiency by avoiding exhaustive pairwise interactions while ensuring spatial and visual coherence between humans and their surroundings. Extensive experiments on monocular video datasets demonstrate that StM significantly outperforms existing state-of-the-art methods in both visual quality and rendering accuracy, particularly at challenging human-scene interaction boundaries.
title Dynamic Avatar-Scene Rendering from Human-centric Context
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.10539