I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vasilakis, Yannis, Bittner, Rachel, Pauwels, Johan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909267718569984
author Vasilakis, Yannis
Bittner, Rachel
Pauwels, Johan
author_facet Vasilakis, Yannis
Bittner, Rachel
Pauwels, Johan
contents Music two-tower multimodal systems integrate audio and text modalities into a joint audio-text space, enabling direct comparison between songs and their corresponding labels. These systems enable new approaches for classification and retrieval, leveraging both modalities. Despite the promising results they have shown for zero-shot classification and retrieval tasks, closer inspection of the embeddings is needed. This paper evaluates the inherent zero-shot properties of joint audio-text spaces for the case-study of instrument recognition. We present an evaluation and analysis of two-tower systems for zero-shot instrument recognition and a detailed analysis of the properties of the pre-joint and joint embeddings spaces. Our findings suggest that audio encoders alone demonstrate good quality, while challenges remain within the text encoder or joint space projection. Specifically, two-tower systems exhibit sensitivity towards specific words, favoring generic prompts over musically informed ones. Despite the large size of textual encoders, they do not yet leverage additional textual context or infer instruments accurately from their descriptions. Lastly, a novel approach for quantifying the semantic meaningfulness of the textual space leveraging an instrument ontology is proposed. This method reveals deficiencies in the systems' understanding of instruments and provides evidence of the need for fine-tuning text encoders on musical data.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18058
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
Vasilakis, Yannis
Bittner, Rachel
Pauwels, Johan
Sound
Computation and Language
Information Retrieval
Machine Learning
Audio and Speech Processing
Music two-tower multimodal systems integrate audio and text modalities into a joint audio-text space, enabling direct comparison between songs and their corresponding labels. These systems enable new approaches for classification and retrieval, leveraging both modalities. Despite the promising results they have shown for zero-shot classification and retrieval tasks, closer inspection of the embeddings is needed. This paper evaluates the inherent zero-shot properties of joint audio-text spaces for the case-study of instrument recognition. We present an evaluation and analysis of two-tower systems for zero-shot instrument recognition and a detailed analysis of the properties of the pre-joint and joint embeddings spaces. Our findings suggest that audio encoders alone demonstrate good quality, while challenges remain within the text encoder or joint space projection. Specifically, two-tower systems exhibit sensitivity towards specific words, favoring generic prompts over musically informed ones. Despite the large size of textual encoders, they do not yet leverage additional textual context or infer instruments accurately from their descriptions. Lastly, a novel approach for quantifying the semantic meaningfulness of the textual space leveraging an instrument ontology is proposed. This method reveals deficiencies in the systems' understanding of instruments and provides evidence of the need for fine-tuning text encoders on musical data.
title I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
topic Sound
Computation and Language
Information Retrieval
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2407.18058