Headset: Human emotion awareness under partial occlusions multimodal dataset

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lohesara, Fatemeh Ghorbani, Freitas, Davi Rabbouni, Guillemot, Christine, Eguiazarian, Karen, Knorr, Sebastian
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909106329092096
author Lohesara, Fatemeh Ghorbani
Freitas, Davi Rabbouni
Guillemot, Christine
Eguiazarian, Karen
Knorr, Sebastian
author_facet Lohesara, Fatemeh Ghorbani
Freitas, Davi Rabbouni
Guillemot, Christine
Eguiazarian, Karen
Knorr, Sebastian
contents The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended Reality (XR) applications, this volumetric data has proven to be an essential technology for future XR elaboration. In this work, we present a new multimodal database to help advance the development of immersive technologies. Our proposed database provides ethically compliant and diverse volumetric data, in particular 27 participants displaying posed facial expressions and subtle body movements while speaking, plus 11 participants wearing head-mounted displays (HMDs). The recording system consists of a volumetric capture (VoCap) studio, including 31 synchronized modules with 62 RGB cameras and 31 depth cameras. In addition to textured meshes, point clouds, and multi-view RGB-D data, we use one Lytro Illum camera for providing light field (LF) data simultaneously. Finally, we also provide an evaluation of our dataset employment with regard to the tasks of facial expression classification, HMDs removal, and point cloud reconstruction. The dataset can be helpful in the evaluation and performance testing of various XR algorithms, including but not limited to facial expression recognition and reconstruction, facial reenactment, and volumetric video. HEADSET and its all associated raw data and license agreement will be publicly available for research purposes.
format Preprint
id arxiv_https___arxiv_org_abs_2402_09107
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Headset: Human emotion awareness under partial occlusions multimodal dataset
Lohesara, Fatemeh Ghorbani
Freitas, Davi Rabbouni
Guillemot, Christine
Eguiazarian, Karen
Knorr, Sebastian
Computer Vision and Pattern Recognition
The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended Reality (XR) applications, this volumetric data has proven to be an essential technology for future XR elaboration. In this work, we present a new multimodal database to help advance the development of immersive technologies. Our proposed database provides ethically compliant and diverse volumetric data, in particular 27 participants displaying posed facial expressions and subtle body movements while speaking, plus 11 participants wearing head-mounted displays (HMDs). The recording system consists of a volumetric capture (VoCap) studio, including 31 synchronized modules with 62 RGB cameras and 31 depth cameras. In addition to textured meshes, point clouds, and multi-view RGB-D data, we use one Lytro Illum camera for providing light field (LF) data simultaneously. Finally, we also provide an evaluation of our dataset employment with regard to the tasks of facial expression classification, HMDs removal, and point cloud reconstruction. The dataset can be helpful in the evaluation and performance testing of various XR algorithms, including but not limited to facial expression recognition and reconstruction, facial reenactment, and volumetric video. HEADSET and its all associated raw data and license agreement will be publicly available for research purposes.
title Headset: Human emotion awareness under partial occlusions multimodal dataset
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.09107