Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bright, Jerrin, Wang, Zhibo, Klepachevskyi, Dmytro, Chen, Yuhao, Rambhatla, Sirisha, Clausi, David, Zelek, John
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914207005409280
author Bright, Jerrin
Wang, Zhibo
Klepachevskyi, Dmytro
Chen, Yuhao
Rambhatla, Sirisha
Clausi, David
Zelek, John
author_facet Bright, Jerrin
Wang, Zhibo
Klepachevskyi, Dmytro
Chen, Yuhao
Rambhatla, Sirisha
Clausi, David
Zelek, John
contents We present Avatar4D, a real-world transferable pipeline for generating customizable synthetic human motion datasets tailored to domain-specific applications. Unlike prior works, which focus on general, everyday motions and offer limited flexibility, our approach provides fine-grained control over body pose, appearance, camera viewpoint, and environmental context, without requiring any manual annotations. To validate the impact of Avatar4D, we focus on sports, where domain-specific human actions and movement patterns pose unique challenges for motion understanding. In this setting, we introduce Syn2Sport, a large-scale synthetic dataset spanning sports, including baseball and ice hockey. Avatar4D features high-fidelity 4D (3D geometry over time) human motion sequences with varying player appearances rendered in diverse environments. We benchmark several state-of-the-art pose estimation models on Syn2Sport and demonstrate their effectiveness for supervised learning, zero-shot transfer to real-world data, and generalization across sports. Furthermore, we evaluate how closely the generated synthetic data aligns with real-world datasets in feature space. Our results highlight the potential of such systems to generate scalable, controllable, and transferable human datasets for diverse domain-specific tasks without relying on domain-specific real data.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16199
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation
Bright, Jerrin
Wang, Zhibo
Klepachevskyi, Dmytro
Chen, Yuhao
Rambhatla, Sirisha
Clausi, David
Zelek, John
Computer Vision and Pattern Recognition
We present Avatar4D, a real-world transferable pipeline for generating customizable synthetic human motion datasets tailored to domain-specific applications. Unlike prior works, which focus on general, everyday motions and offer limited flexibility, our approach provides fine-grained control over body pose, appearance, camera viewpoint, and environmental context, without requiring any manual annotations. To validate the impact of Avatar4D, we focus on sports, where domain-specific human actions and movement patterns pose unique challenges for motion understanding. In this setting, we introduce Syn2Sport, a large-scale synthetic dataset spanning sports, including baseball and ice hockey. Avatar4D features high-fidelity 4D (3D geometry over time) human motion sequences with varying player appearances rendered in diverse environments. We benchmark several state-of-the-art pose estimation models on Syn2Sport and demonstrate their effectiveness for supervised learning, zero-shot transfer to real-world data, and generalization across sports. Furthermore, we evaluate how closely the generated synthetic data aligns with real-world datasets in feature space. Our results highlight the potential of such systems to generate scalable, controllable, and transferable human datasets for diverse domain-specific tasks without relying on domain-specific real data.
title Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.16199