Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hu, Xinhao, Zhang, Yiyi, Zhang, Liqing, Zhang, Jianfu
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911675428372480
author Hu, Xinhao
Zhang, Yiyi
Zhang, Liqing
Zhang, Jianfu
author_facet Hu, Xinhao
Zhang, Yiyi
Zhang, Liqing
Zhang, Jianfu
contents Pedestrian motion, due to its causal nature, is strongly influenced by domain gaps arising from discrepancies between training and testing data distributions. Focusing on 3D human pose estimation, this work presents a controllable human pose generation framework that synthesizes diverse video data by systematically varying poses, backgrounds, and camera viewpoints. This generative augmentation enriches training datasets, enhances model generalization, and alleviates the limitations of existing methods in handling domain discrepancies. By leveraging both indoor/real-world and outdoor/virtual datasets, we perform cross-domain data fusion and controllable video generation to construct enriched training data, tailored to realistic deployment settings. Extensive experiments show that the augmented datasets significantly improve model performance on unseen scenarios and datasets, validating the effectiveness of the proposed approach.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12198
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
Hu, Xinhao
Zhang, Yiyi
Zhang, Liqing
Zhang, Jianfu
Computer Vision and Pattern Recognition
Pedestrian motion, due to its causal nature, is strongly influenced by domain gaps arising from discrepancies between training and testing data distributions. Focusing on 3D human pose estimation, this work presents a controllable human pose generation framework that synthesizes diverse video data by systematically varying poses, backgrounds, and camera viewpoints. This generative augmentation enriches training datasets, enhances model generalization, and alleviates the limitations of existing methods in handling domain discrepancies. By leveraging both indoor/real-world and outdoor/virtual datasets, we perform cross-domain data fusion and controllable video generation to construct enriched training data, tailored to realistic deployment settings. Extensive experiments show that the augmented datasets significantly improve model performance on unseen scenarios and datasets, validating the effectiveness of the proposed approach.
title Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.12198