Metric Learning with Progressive Self-Distillation for Audio-Visual Embedding Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zeng, Donghuo, Ikeda, Kazushi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917893892997120
author Zeng, Donghuo
Ikeda, Kazushi
author_facet Zeng, Donghuo
Ikeda, Kazushi
contents Metric learning projects samples into an embedded space, where similarities and dissimilarities are quantified based on their learned representations. However, existing methods often rely on label-guided representation learning, where representations of different modalities, such as audio and visual data, are aligned based on annotated labels. This approach tends to underutilize latent complex features and potential relationships inherent in the distributions of audio and visual data that are not directly tied to the labels, resulting in suboptimal performance in audio-visual embedding learning. To address this issue, we propose a novel architecture that integrates cross-modal triplet loss with progressive self-distillation. Our method enhances representation learning by leveraging inherent distributions and dynamically refining soft audio-visual alignments -- probabilistic alignments between audio and visual data that capture the inherent relationships beyond explicit labels. Specifically, the model distills audio-visual distribution-based knowledge from annotated labels in a subset of each batch. This self-distilled knowledge is used t
format Preprint
id arxiv_https___arxiv_org_abs_2501_09608
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Metric Learning with Progressive Self-Distillation for Audio-Visual Embedding Learning
Zeng, Donghuo
Ikeda, Kazushi
Sound
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
Audio and Speech Processing
Metric learning projects samples into an embedded space, where similarities and dissimilarities are quantified based on their learned representations. However, existing methods often rely on label-guided representation learning, where representations of different modalities, such as audio and visual data, are aligned based on annotated labels. This approach tends to underutilize latent complex features and potential relationships inherent in the distributions of audio and visual data that are not directly tied to the labels, resulting in suboptimal performance in audio-visual embedding learning. To address this issue, we propose a novel architecture that integrates cross-modal triplet loss with progressive self-distillation. Our method enhances representation learning by leveraging inherent distributions and dynamically refining soft audio-visual alignments -- probabilistic alignments between audio and visual data that capture the inherent relationships beyond explicit labels. Specifically, the model distills audio-visual distribution-based knowledge from annotated labels in a subset of each batch. This self-distilled knowledge is used t
title Metric Learning with Progressive Self-Distillation for Audio-Visual Embedding Learning
topic Sound
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2501.09608