Enhancing Human Pose Estimation in Ancient Vase Paintings via Perceptually-grounded Style Transfer Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Madhu, Prathmesh, Villar-Corrales, Angel, Kosti, Ronak, Bendschus, Torsten, Reinhardt, Corinna, Bell, Peter, Maier, Andreas, Christlein, Vincent
Formato: Preprint
Publicado: 2020
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929255164674048
author Madhu, Prathmesh
Villar-Corrales, Angel
Kosti, Ronak
Bendschus, Torsten
Reinhardt, Corinna
Bell, Peter
Maier, Andreas
Christlein, Vincent
author_facet Madhu, Prathmesh
Villar-Corrales, Angel
Kosti, Ronak
Bendschus, Torsten
Reinhardt, Corinna
Bell, Peter
Maier, Andreas
Christlein, Vincent
contents Human pose estimation (HPE) is a central part of understanding the visual narration and body movements of characters depicted in artwork collections, such as Greek vase paintings. Unfortunately, existing HPE methods do not generalise well across domains resulting in poorly recognized poses. Therefore, we propose a two step approach: (1) adapting a dataset of natural images of known person and pose annotations to the style of Greek vase paintings by means of image style-transfer. We introduce a perceptually-grounded style transfer training to enforce perceptual consistency. Then, we fine-tune the base model with this newly created dataset. We show that using style-transfer learning significantly improves the SOTA performance on unlabelled data by more than 6% mean average precision (mAP) as well as mean average recall (mAR). (2) To improve the already strong results further, we created a small dataset (ClassArch) consisting of ancient Greek vase paintings from the 6-5th century BCE with person and pose annotations. We show that fine-tuning on this data with a style-transferred model improves the performance further. In a thorough ablation study, we give a targeted analysis of the influence of style intensities, revealing that the model learns generic domain styles. Additionally, we provide a pose-based image retrieval to demonstrate the effectiveness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2012_05616
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Enhancing Human Pose Estimation in Ancient Vase Paintings via Perceptually-grounded Style Transfer Learning
Madhu, Prathmesh
Villar-Corrales, Angel
Kosti, Ronak
Bendschus, Torsten
Reinhardt, Corinna
Bell, Peter
Maier, Andreas
Christlein, Vincent
Computer Vision and Pattern Recognition
Human pose estimation (HPE) is a central part of understanding the visual narration and body movements of characters depicted in artwork collections, such as Greek vase paintings. Unfortunately, existing HPE methods do not generalise well across domains resulting in poorly recognized poses. Therefore, we propose a two step approach: (1) adapting a dataset of natural images of known person and pose annotations to the style of Greek vase paintings by means of image style-transfer. We introduce a perceptually-grounded style transfer training to enforce perceptual consistency. Then, we fine-tune the base model with this newly created dataset. We show that using style-transfer learning significantly improves the SOTA performance on unlabelled data by more than 6% mean average precision (mAP) as well as mean average recall (mAR). (2) To improve the already strong results further, we created a small dataset (ClassArch) consisting of ancient Greek vase paintings from the 6-5th century BCE with person and pose annotations. We show that fine-tuning on this data with a style-transferred model improves the performance further. In a thorough ablation study, we give a targeted analysis of the influence of style intensities, revealing that the model learns generic domain styles. Additionally, we provide a pose-based image retrieval to demonstrate the effectiveness of our method.
title Enhancing Human Pose Estimation in Ancient Vase Paintings via Perceptually-grounded Style Transfer Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2012.05616