Saved in:
| Main Author: | |
|---|---|
| Format: | Recurso digital |
| Language: | |
| Published: |
Zenodo
2023
|
| Subjects: | |
| Online Access: | https://doi.org/10.5281/zenodo.8385275 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Table of Contents:
- <p>Freesound is an online platform where people using sounds for various purposes can share or download audio clips. In such platforms, it is crucial that the users are provided with accurate sound recommendations, which becomes challenging due to the large size of the audio collection, complexity of the sound properties, and the human aspect of the recommendations. To provide sound recommendations, Freesound features a "similar sounds" function. However, this function primarily relies on creating a digital representation of audio clips that assesses the acoustic characteristics of sounds, which proves to be insufficient for accurately capturing their semantic properties. This limitation reduces the content-based retrieval capa-bilities of Freesound users. Moreover, the audio representation is created by hand-picking features that were engineered using domain knowledge. Today, in various fields related to audio, this approach has been replaced by using neural networks as feature extractors. In this work, we search for pretrained general-purpose neural net-works that can be used to represent the semantic content of audio clips. We choose 8 such models and compare their semantic sound similarity performances both ob-jectively and subjectively. During the integration of deep embeddings in the sound similarity system, we explore numerous design choices and share valuable insights. We use the FSD50K evaluation set for all experiments and report various objective metrics using the sound class hierarchy to perform multi-level analysis, including class- and family-level. We find out that most of the neural networks outperform the hand-made representation subjectively and objectively. Specifically, the multi-modal representation learning model CLAP that uses natural language and audio as modalities outperforms other models by a significant margin, while the models that attempt to leverage the CLIP model for creating tri-modal representations fail.</p>