Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Clarke, Jason, Gotoh, Yoshihiko, Goetze, Stefan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918067805618176
author Clarke, Jason
Gotoh, Yoshihiko
Goetze, Stefan
author_facet Clarke, Jason
Gotoh, Yoshihiko
Goetze, Stefan
contents Audiovisual active speaker detection (ASD) is conventionally performed by modelling the temporal synchronisation of acoustic and visual speech cues. In egocentric recordings, however, the efficacy of synchronisation-based methods is compromised by occlusions, motion blur, and adverse acoustic conditions. In this work, a novel framework is proposed that exclusively leverages cross-modal face-voice associations to determine speaker activity. An existing face-voice association model is integrated with a transformer-based encoder that aggregates facial identity information by dynamically weighting each frame based on its visual quality. This system is then coupled with a front-end utterance segmentation method, producing a complete ASD system. This work demonstrates that the proposed system, Self-Lifting for audiovisual active speaker detection(SL-ASD), achieves performance comparable to, and in certain cases exceeding, that of parameter-intensive synchronisation-based approaches with significantly fewer learnable parameters, thereby validating the feasibility of substituting strict audiovisual synchronisation modelling with flexible biometric associations in challenging egocentric scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
Clarke, Jason
Gotoh, Yoshihiko
Goetze, Stefan
Multimedia
Sound
Audio and Speech Processing
Audiovisual active speaker detection (ASD) is conventionally performed by modelling the temporal synchronisation of acoustic and visual speech cues. In egocentric recordings, however, the efficacy of synchronisation-based methods is compromised by occlusions, motion blur, and adverse acoustic conditions. In this work, a novel framework is proposed that exclusively leverages cross-modal face-voice associations to determine speaker activity. An existing face-voice association model is integrated with a transformer-based encoder that aggregates facial identity information by dynamically weighting each frame based on its visual quality. This system is then coupled with a front-end utterance segmentation method, producing a complete ASD system. This work demonstrates that the proposed system, Self-Lifting for audiovisual active speaker detection(SL-ASD), achieves performance comparable to, and in certain cases exceeding, that of parameter-intensive synchronisation-based approaches with significantly fewer learnable parameters, thereby validating the feasibility of substituting strict audiovisual synchronisation modelling with flexible biometric associations in challenging egocentric scenarios.
title Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
topic Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.18055