PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hannan, Abdul, Manzoor, Muhammad Arslan, Nawaz, Shah, Liaqat, Muhammad Irzam, Schedl, Markus, Noman, Mubashir
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918037107507200
author Hannan, Abdul
Manzoor, Muhammad Arslan
Nawaz, Shah
Liaqat, Muhammad Irzam
Schedl, Markus
Noman, Mubashir
author_facet Hannan, Abdul
Manzoor, Muhammad Arslan
Nawaz, Shah
Liaqat, Muhammad Irzam
Schedl, Markus
Noman, Mubashir
contents We study the task of learning association between faces and voices, which is gaining interest in the multimodal community lately. These methods suffer from the deliberate crafting of negative mining procedures as well as the reliance on the distant margin parameter. These issues are addressed by learning a joint embedding space in which orthogonality constraints are applied to the fused embeddings of faces and voices. However, embedding spaces of faces and voices possess different characteristics and require spaces to be aligned before fusing them. To this end, we propose a method that accurately aligns the embedding spaces and fuses them with an enhanced gated fusion thereby improving the performance of face-voice association. Extensive experiments on the VoxCeleb dataset reveals the merits of the proposed approach.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17002
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association
Hannan, Abdul
Manzoor, Muhammad Arslan
Nawaz, Shah
Liaqat, Muhammad Irzam
Schedl, Markus
Noman, Mubashir
Computer Vision and Pattern Recognition
Artificial Intelligence
We study the task of learning association between faces and voices, which is gaining interest in the multimodal community lately. These methods suffer from the deliberate crafting of negative mining procedures as well as the reliance on the distant margin parameter. These issues are addressed by learning a joint embedding space in which orthogonality constraints are applied to the fused embeddings of faces and voices. However, embedding spaces of faces and voices possess different characteristics and require spaces to be aligned before fusing them. To this end, we propose a method that accurately aligns the embedding spaces and fuses them with an enhanced gated fusion thereby improving the performance of face-voice association. Extensive experiments on the VoxCeleb dataset reveals the merits of the proposed approach.
title PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.17002