Discogs-VI: A Musical Version Identification Dataset Based on Public Editorial Metadata

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Araz, R. Oguz, Serra, Xavier, Bogdanov, Dmitry
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914985260613632
author Araz, R. Oguz
Serra, Xavier
Bogdanov, Dmitry
author_facet Araz, R. Oguz
Serra, Xavier
Bogdanov, Dmitry
contents Current version identification (VI) datasets often lack sufficient size and musical diversity to train robust neural networks (NNs). Additionally, their non-representative clique size distributions prevent realistic system evaluations. To address these challenges, we explore the untapped potential of the rich editorial metadata in the Discogs music database and create a large dataset of musical versions containing about 1,900,000 versions across 348,000 cliques. Utilizing a high-precision search algorithm, we map this dataset to official music uploads on YouTube, resulting in a dataset of approximately 493,000 versions across 98,000 cliques. This dataset offers over nine times the number of cliques and over four times the number of versions than existing datasets. We demonstrate the utility of our dataset by training a baseline NN without extensive model complexities or data augmentations, which achieves competitive results on the SHS100K and Da-TACOS datasets. Our dataset, along with the tools used for its creation, the extracted audio features, and a trained model, are all publicly available online.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17400
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Discogs-VI: A Musical Version Identification Dataset Based on Public Editorial Metadata
Araz, R. Oguz
Serra, Xavier
Bogdanov, Dmitry
Sound
Audio and Speech Processing
Current version identification (VI) datasets often lack sufficient size and musical diversity to train robust neural networks (NNs). Additionally, their non-representative clique size distributions prevent realistic system evaluations. To address these challenges, we explore the untapped potential of the rich editorial metadata in the Discogs music database and create a large dataset of musical versions containing about 1,900,000 versions across 348,000 cliques. Utilizing a high-precision search algorithm, we map this dataset to official music uploads on YouTube, resulting in a dataset of approximately 493,000 versions across 98,000 cliques. This dataset offers over nine times the number of cliques and over four times the number of versions than existing datasets. We demonstrate the utility of our dataset by training a baseline NN without extensive model complexities or data augmentations, which achieves competitive results on the SHS100K and Da-TACOS datasets. Our dataset, along with the tools used for its creation, the extracted audio features, and a trained model, are all publicly available online.
title Discogs-VI: A Musical Version Identification Dataset Based on Public Editorial Metadata
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.17400