Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866911675428372480 |
|---|---|
| author | Hu, Xinhao Zhang, Yiyi Zhang, Liqing Zhang, Jianfu |
| author_facet | Hu, Xinhao Zhang, Yiyi Zhang, Liqing Zhang, Jianfu |
| contents | Pedestrian motion, due to its causal nature, is strongly influenced by domain gaps arising from discrepancies between training and testing data distributions. Focusing on 3D human pose estimation, this work presents a controllable human pose generation framework that synthesizes diverse video data by systematically varying poses, backgrounds, and camera viewpoints. This generative augmentation enriches training datasets, enhances model generalization, and alleviates the limitations of existing methods in handling domain discrepancies. By leveraging both indoor/real-world and outdoor/virtual datasets, we perform cross-domain data fusion and controllable video generation to construct enriched training data, tailored to realistic deployment settings. Extensive experiments show that the augmented datasets significantly improve model performance on unseen scenarios and datasets, validating the effectiveness of the proposed approach. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_12198 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation Hu, Xinhao Zhang, Yiyi Zhang, Liqing Zhang, Jianfu Computer Vision and Pattern Recognition Pedestrian motion, due to its causal nature, is strongly influenced by domain gaps arising from discrepancies between training and testing data distributions. Focusing on 3D human pose estimation, this work presents a controllable human pose generation framework that synthesizes diverse video data by systematically varying poses, backgrounds, and camera viewpoints. This generative augmentation enriches training datasets, enhances model generalization, and alleviates the limitations of existing methods in handling domain discrepancies. By leveraging both indoor/real-world and outdoor/virtual datasets, we perform cross-domain data fusion and controllable video generation to construct enriched training data, tailored to realistic deployment settings. Extensive experiments show that the augmented datasets significantly improve model performance on unseen scenarios and datasets, validating the effectiveness of the proposed approach. |
| title | Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.12198 |