Diversifying Human Pose in Synthetic Data for Aerial-view Human Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shen, Yi-Ting, Lee, Hyungtae, Kwon, Heesung, Bhattacharyya, Shuvra S.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913890517909504
author Shen, Yi-Ting
Lee, Hyungtae
Kwon, Heesung
Bhattacharyya, Shuvra S.
author_facet Shen, Yi-Ting
Lee, Hyungtae
Kwon, Heesung
Bhattacharyya, Shuvra S.
contents Synthetic data generation has emerged as a promising solution to the data scarcity issue in aerial-view human detection. However, creating datasets that accurately reflect varying real-world human appearances, particularly diverse poses, remains challenging and labor-intensive. To address this, we propose SynPoseDiv, a novel framework that diversifies human poses within existing synthetic datasets. SynPoseDiv tackles two key challenges: generating realistic, diverse 3D human poses using a diffusion-based pose generator, and producing images of virtual characters in novel poses through a source-to-target image translator. The framework incrementally transitions characters into new poses using optimized pose sequences identified via Dijkstra's algorithm. Experiments demonstrate that SynPoseDiv significantly improves detection accuracy across multiple aerial-view human detection benchmarks, especially in low-shot scenarios, and remains effective regardless of the training approach or dataset size.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15939
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diversifying Human Pose in Synthetic Data for Aerial-view Human Detection
Shen, Yi-Ting
Lee, Hyungtae
Kwon, Heesung
Bhattacharyya, Shuvra S.
Computer Vision and Pattern Recognition
Synthetic data generation has emerged as a promising solution to the data scarcity issue in aerial-view human detection. However, creating datasets that accurately reflect varying real-world human appearances, particularly diverse poses, remains challenging and labor-intensive. To address this, we propose SynPoseDiv, a novel framework that diversifies human poses within existing synthetic datasets. SynPoseDiv tackles two key challenges: generating realistic, diverse 3D human poses using a diffusion-based pose generator, and producing images of virtual characters in novel poses through a source-to-target image translator. The framework incrementally transitions characters into new poses using optimized pose sequences identified via Dijkstra's algorithm. Experiments demonstrate that SynPoseDiv significantly improves detection accuracy across multiple aerial-view human detection benchmarks, especially in low-shot scenarios, and remains effective regardless of the training approach or dataset size.
title Diversifying Human Pose in Synthetic Data for Aerial-view Human Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.15939