Improving the Robustness of 3D Human Pose Estimation: A Benchmark and Learning from Noisy Input

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hoang, Trung-Hieu, Zehni, Mona, Phan, Huy, Vo, Duc Minh, Do, Minh N.
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929314819211264
author Hoang, Trung-Hieu
Zehni, Mona
Phan, Huy
Vo, Duc Minh
Do, Minh N.
author_facet Hoang, Trung-Hieu
Zehni, Mona
Phan, Huy
Vo, Duc Minh
Do, Minh N.
contents Despite the promising performance of current 3D human pose estimation techniques, understanding and enhancing their generalization on challenging in-the-wild videos remain an open problem. In this work, we focus on the robustness of 2D-to-3D pose lifters. To this end, we develop two benchmark datasets, namely Human3.6M-C and HumanEva-I-C, to examine the robustness of video-based 3D pose lifters to a wide range of common video corruptions including temporary occlusion, motion blur, and pixel-level noise. We observe the poor generalization of state-of-the-art 3D pose lifters in the presence of corruption and establish two techniques to tackle this issue. First, we introduce Temporal Additive Gaussian Noise (TAGN) as a simple yet effective 2D input pose data augmentation. Additionally, to incorporate the confidence scores output by the 2D pose detectors, we design a confidence-aware convolution (CA-Conv) block. Extensively tested on corrupted videos, the proposed strategies consistently boost the robustness of 3D pose lifters and serve as new baselines for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2312_06797
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Improving the Robustness of 3D Human Pose Estimation: A Benchmark and Learning from Noisy Input
Hoang, Trung-Hieu
Zehni, Mona
Phan, Huy
Vo, Duc Minh
Do, Minh N.
Computer Vision and Pattern Recognition
Despite the promising performance of current 3D human pose estimation techniques, understanding and enhancing their generalization on challenging in-the-wild videos remain an open problem. In this work, we focus on the robustness of 2D-to-3D pose lifters. To this end, we develop two benchmark datasets, namely Human3.6M-C and HumanEva-I-C, to examine the robustness of video-based 3D pose lifters to a wide range of common video corruptions including temporary occlusion, motion blur, and pixel-level noise. We observe the poor generalization of state-of-the-art 3D pose lifters in the presence of corruption and establish two techniques to tackle this issue. First, we introduce Temporal Additive Gaussian Noise (TAGN) as a simple yet effective 2D input pose data augmentation. Additionally, to incorporate the confidence scores output by the 2D pose detectors, we design a confidence-aware convolution (CA-Conv) block. Extensively tested on corrupted videos, the proposed strategies consistently boost the robustness of 3D pose lifters and serve as new baselines for future research.
title Improving the Robustness of 3D Human Pose Estimation: A Benchmark and Learning from Noisy Input
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.06797