SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rajasekhar, Gnana Praveen, Alam, Jahangir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909658143260672
author Rajasekhar, Gnana Praveen
Alam, Jahangir
author_facet Rajasekhar, Gnana Praveen
Alam, Jahangir
contents Conventional audio-visual methods for speaker verification rely on large amounts of labeled data and separate modality-specific architectures, which is computationally expensive, limiting their scalability. To address these problems, we propose a self-supervised learning framework based on contrastive learning with asymmetric masking and masked data modeling to obtain robust audiovisual feature representations. In particular, we employ a unified framework for self-supervised audiovisual speaker verification using a single shared backbone for audio and visual inputs, leveraging the versatility of vision transformers. The proposed unified framework can handle audio, visual, or audiovisual inputs using a single shared vision transformer backbone during training and testing while being computationally efficient and robust to missing modalities. Extensive experiments demonstrate that our method achieves competitive performance without labeled data while reducing computational costs compared to traditional approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
Rajasekhar, Gnana Praveen
Alam, Jahangir
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
Conventional audio-visual methods for speaker verification rely on large amounts of labeled data and separate modality-specific architectures, which is computationally expensive, limiting their scalability. To address these problems, we propose a self-supervised learning framework based on contrastive learning with asymmetric masking and masked data modeling to obtain robust audiovisual feature representations. In particular, we employ a unified framework for self-supervised audiovisual speaker verification using a single shared backbone for audio and visual inputs, leveraging the versatility of vision transformers. The proposed unified framework can handle audio, visual, or audiovisual inputs using a single shared vision transformer backbone during training and testing while being computationally efficient and robust to missing modalities. Extensive experiments demonstrate that our method achieves competitive performance without labeled data while reducing computational costs compared to traditional approaches.
title SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
topic Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.17694