Shared Multi-modal Embedding Space for Face-Voice Association
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911302032556032 |
|---|---|
| author | Simic, Christopher Riedhammer, Korbinian Bocklet, Tobias |
| author_facet | Simic, Christopher Riedhammer, Korbinian Bocklet, Tobias |
| contents | The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model was not trained. Our approach consists of separate uni-modal processing pipelines with general face and voice feature extraction, complemented by additional age-gender feature extraction to support prediction. The resulting single-modal features are projected into a shared embedding space and trained with an Adaptive Angular Margin (AAM) loss. Our approach achieved first place in the FAME 2026 challenge, with an average Equal-Error Rate (EER) of 23.99%. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_04814 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Shared Multi-modal Embedding Space for Face-Voice Association Simic, Christopher Riedhammer, Korbinian Bocklet, Tobias Sound Computer Vision and Pattern Recognition The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model was not trained. Our approach consists of separate uni-modal processing pipelines with general face and voice feature extraction, complemented by additional age-gender feature extraction to support prediction. The resulting single-modal features are projected into a shared embedding space and trained with an Adaptive Angular Margin (AAM) loss. Our approach achieved first place in the FAME 2026 challenge, with an average Equal-Error Rate (EER) of 23.99%. |
| title | Shared Multi-modal Embedding Space for Face-Voice Association |
| topic | Sound Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.04814 |