Shared Multi-modal Embedding Space for Face-Voice Association

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Simic, Christopher, Riedhammer, Korbinian, Bocklet, Tobias
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911302032556032
author Simic, Christopher
Riedhammer, Korbinian
Bocklet, Tobias
author_facet Simic, Christopher
Riedhammer, Korbinian
Bocklet, Tobias
contents The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model was not trained. Our approach consists of separate uni-modal processing pipelines with general face and voice feature extraction, complemented by additional age-gender feature extraction to support prediction. The resulting single-modal features are projected into a shared embedding space and trained with an Adaptive Angular Margin (AAM) loss. Our approach achieved first place in the FAME 2026 challenge, with an average Equal-Error Rate (EER) of 23.99%.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04814
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Shared Multi-modal Embedding Space for Face-Voice Association
Simic, Christopher
Riedhammer, Korbinian
Bocklet, Tobias
Sound
Computer Vision and Pattern Recognition
The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model was not trained. Our approach consists of separate uni-modal processing pipelines with general face and voice feature extraction, complemented by additional age-gender feature extraction to support prediction. The resulting single-modal features are projected into a shared embedding space and trained with an Adaptive Angular Margin (AAM) loss. Our approach achieved first place in the FAME 2026 challenge, with an average Equal-Error Rate (EER) of 23.99%.
title Shared Multi-modal Embedding Space for Face-Voice Association
topic Sound
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.04814