Neuromorphic Facial Analysis with Cross-Modal Supervision

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Becattini, Federico, Cultrera, Luca, Berlincioni, Lorenzo, Ferrari, Claudio, Leonardo, Andrea, Del Bimbo, Alberto
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909317375983616
author Becattini, Federico
Cultrera, Luca
Berlincioni, Lorenzo
Ferrari, Claudio
Leonardo, Andrea
Del Bimbo, Alberto
author_facet Becattini, Federico
Cultrera, Luca
Berlincioni, Lorenzo
Ferrari, Claudio
Leonardo, Andrea
Del Bimbo, Alberto
contents Traditional approaches for analyzing RGB frames are capable of providing a fine-grained understanding of a face from different angles by inferring emotions, poses, shapes, landmarks. However, when it comes to subtle movements standard RGB cameras might fall behind due to their latency, making it hard to detect micro-movements that carry highly informative cues to infer the true emotions of a subject. To address this issue, the usage of event cameras to analyze faces is gaining increasing interest. Nonetheless, all the expertise matured for RGB processing is not directly transferrable to neuromorphic data due to a strong domain shift and intrinsic differences in how data is represented. The lack of labeled data can be considered one of the main causes of this gap, yet gathering data is harder in the event domain since it cannot be crawled from the web and labeling frames should take into account event aggregation rates and the fact that static parts might not be visible in certain frames. In this paper, we first present FACEMORPHIC, a multimodal temporally synchronized face dataset comprising both RGB videos and event streams. The data is labeled at a video level with facial Action Units and also contains streams collected with a variety of applications in mind, ranging from 3D shape estimation to lip-reading. We then show how temporal synchronization can allow effective neuromorphic face analysis without the need to manually annotate videos: we instead leverage cross-modal supervision bridging the domain gap by representing face shapes in a 3D space.
format Preprint
id arxiv_https___arxiv_org_abs_2409_10213
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Neuromorphic Facial Analysis with Cross-Modal Supervision
Becattini, Federico
Cultrera, Luca
Berlincioni, Lorenzo
Ferrari, Claudio
Leonardo, Andrea
Del Bimbo, Alberto
Computer Vision and Pattern Recognition
Traditional approaches for analyzing RGB frames are capable of providing a fine-grained understanding of a face from different angles by inferring emotions, poses, shapes, landmarks. However, when it comes to subtle movements standard RGB cameras might fall behind due to their latency, making it hard to detect micro-movements that carry highly informative cues to infer the true emotions of a subject. To address this issue, the usage of event cameras to analyze faces is gaining increasing interest. Nonetheless, all the expertise matured for RGB processing is not directly transferrable to neuromorphic data due to a strong domain shift and intrinsic differences in how data is represented. The lack of labeled data can be considered one of the main causes of this gap, yet gathering data is harder in the event domain since it cannot be crawled from the web and labeling frames should take into account event aggregation rates and the fact that static parts might not be visible in certain frames. In this paper, we first present FACEMORPHIC, a multimodal temporally synchronized face dataset comprising both RGB videos and event streams. The data is labeled at a video level with facial Action Units and also contains streams collected with a variety of applications in mind, ranging from 3D shape estimation to lip-reading. We then show how temporal synchronization can allow effective neuromorphic face analysis without the need to manually annotate videos: we instead leverage cross-modal supervision bridging the domain gap by representing face shapes in a 3D space.
title Neuromorphic Facial Analysis with Cross-Modal Supervision
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.10213