Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914207005409280 |
|---|---|
| author | Bright, Jerrin Wang, Zhibo Klepachevskyi, Dmytro Chen, Yuhao Rambhatla, Sirisha Clausi, David Zelek, John |
| author_facet | Bright, Jerrin Wang, Zhibo Klepachevskyi, Dmytro Chen, Yuhao Rambhatla, Sirisha Clausi, David Zelek, John |
| contents | We present Avatar4D, a real-world transferable pipeline for generating customizable synthetic human motion datasets tailored to domain-specific applications. Unlike prior works, which focus on general, everyday motions and offer limited flexibility, our approach provides fine-grained control over body pose, appearance, camera viewpoint, and environmental context, without requiring any manual annotations. To validate the impact of Avatar4D, we focus on sports, where domain-specific human actions and movement patterns pose unique challenges for motion understanding. In this setting, we introduce Syn2Sport, a large-scale synthetic dataset spanning sports, including baseball and ice hockey. Avatar4D features high-fidelity 4D (3D geometry over time) human motion sequences with varying player appearances rendered in diverse environments. We benchmark several state-of-the-art pose estimation models on Syn2Sport and demonstrate their effectiveness for supervised learning, zero-shot transfer to real-world data, and generalization across sports. Furthermore, we evaluate how closely the generated synthetic data aligns with real-world datasets in feature space. Our results highlight the potential of such systems to generate scalable, controllable, and transferable human datasets for diverse domain-specific tasks without relying on domain-specific real data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_16199 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation Bright, Jerrin Wang, Zhibo Klepachevskyi, Dmytro Chen, Yuhao Rambhatla, Sirisha Clausi, David Zelek, John Computer Vision and Pattern Recognition We present Avatar4D, a real-world transferable pipeline for generating customizable synthetic human motion datasets tailored to domain-specific applications. Unlike prior works, which focus on general, everyday motions and offer limited flexibility, our approach provides fine-grained control over body pose, appearance, camera viewpoint, and environmental context, without requiring any manual annotations. To validate the impact of Avatar4D, we focus on sports, where domain-specific human actions and movement patterns pose unique challenges for motion understanding. In this setting, we introduce Syn2Sport, a large-scale synthetic dataset spanning sports, including baseball and ice hockey. Avatar4D features high-fidelity 4D (3D geometry over time) human motion sequences with varying player appearances rendered in diverse environments. We benchmark several state-of-the-art pose estimation models on Syn2Sport and demonstrate their effectiveness for supervised learning, zero-shot transfer to real-world data, and generalization across sports. Furthermore, we evaluate how closely the generated synthetic data aligns with real-world datasets in feature space. Our results highlight the potential of such systems to generate scalable, controllable, and transferable human datasets for diverse domain-specific tasks without relying on domain-specific real data. |
| title | Avatar4D: Synthesizing Domain-Specific 4D Humans for Real-World Pose Estimation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.16199 |