Diversifying Human Pose in Synthetic Data for Aerial-view Human Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866913890517909504 |
|---|---|
| author | Shen, Yi-Ting Lee, Hyungtae Kwon, Heesung Bhattacharyya, Shuvra S. |
| author_facet | Shen, Yi-Ting Lee, Hyungtae Kwon, Heesung Bhattacharyya, Shuvra S. |
| contents | Synthetic data generation has emerged as a promising solution to the data scarcity issue in aerial-view human detection. However, creating datasets that accurately reflect varying real-world human appearances, particularly diverse poses, remains challenging and labor-intensive. To address this, we propose SynPoseDiv, a novel framework that diversifies human poses within existing synthetic datasets. SynPoseDiv tackles two key challenges: generating realistic, diverse 3D human poses using a diffusion-based pose generator, and producing images of virtual characters in novel poses through a source-to-target image translator. The framework incrementally transitions characters into new poses using optimized pose sequences identified via Dijkstra's algorithm. Experiments demonstrate that SynPoseDiv significantly improves detection accuracy across multiple aerial-view human detection benchmarks, especially in low-shot scenarios, and remains effective regardless of the training approach or dataset size. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_15939 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Diversifying Human Pose in Synthetic Data for Aerial-view Human Detection Shen, Yi-Ting Lee, Hyungtae Kwon, Heesung Bhattacharyya, Shuvra S. Computer Vision and Pattern Recognition Synthetic data generation has emerged as a promising solution to the data scarcity issue in aerial-view human detection. However, creating datasets that accurately reflect varying real-world human appearances, particularly diverse poses, remains challenging and labor-intensive. To address this, we propose SynPoseDiv, a novel framework that diversifies human poses within existing synthetic datasets. SynPoseDiv tackles two key challenges: generating realistic, diverse 3D human poses using a diffusion-based pose generator, and producing images of virtual characters in novel poses through a source-to-target image translator. The framework incrementally transitions characters into new poses using optimized pose sequences identified via Dijkstra's algorithm. Experiments demonstrate that SynPoseDiv significantly improves detection accuracy across multiple aerial-view human detection benchmarks, especially in low-shot scenarios, and remains effective regardless of the training approach or dataset size. |
| title | Diversifying Human Pose in Synthetic Data for Aerial-view Human Detection |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2405.15939 |