The Spheres Dataset: Multitrack Orchestral Recordings for Music Source Separation and Information Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Garcia-Martinez, Jaime, Diaz-Guerra, David, Anderson, John, Falcon-Perez, Ricardo, Cabañas-Molero, Pablo, Virtanen, Tuomas, Carabias-Orti, Julio J., Vera-Candeas, Pedro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917492666925056
author Garcia-Martinez, Jaime
Diaz-Guerra, David
Anderson, John
Falcon-Perez, Ricardo
Cabañas-Molero, Pablo
Virtanen, Tuomas
Carabias-Orti, Julio J.
Vera-Candeas, Pedro
author_facet Garcia-Martinez, Jaime
Diaz-Guerra, David
Anderson, John
Falcon-Perez, Ricardo
Cabañas-Molero, Pablo
Virtanen, Tuomas
Carabias-Orti, Julio J.
Vera-Candeas, Pedro
contents This paper introduces The Spheres dataset, multitrack orchestral recordings designed to advance machine learning research in music source separation and related MIR tasks within the classical music domain. The dataset is composed of over one hour recordings of musical pieces performed by the Colibrì Ensemble at The Spheres recording studio, capturing two canonical works - Tchaikovsky's Romeo and Juliet and Mozart's Symphony No. 40 - along with chromatic scales and solo excerpts for each instrument. The recording setup employed 23 microphones, including close spot, main, and ambient microphones, enabling the creation of realistic stereo mixes with controlled bleeding and providing isolated stems for supervised training of source separation models. In addition, room impulse responses were estimated for each instrument position, offering valuable acoustic characterization of the recording space. We present the dataset structure, acoustic analysis, and baseline evaluations using X-UMX based models for orchestral family separation and microphone debleeding. Results highlight both the potential and the challenges of source separation in complex orchestral scenarios, underscoring the dataset's value for benchmarking and for exploring new approaches to separation, localization, dereverberation, and immersive rendering of classical music.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21247
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Spheres Dataset: Multitrack Orchestral Recordings for Music Source Separation and Information Retrieval
Garcia-Martinez, Jaime
Diaz-Guerra, David
Anderson, John
Falcon-Perez, Ricardo
Cabañas-Molero, Pablo
Virtanen, Tuomas
Carabias-Orti, Julio J.
Vera-Candeas, Pedro
Audio and Speech Processing
Machine Learning
Sound
This paper introduces The Spheres dataset, multitrack orchestral recordings designed to advance machine learning research in music source separation and related MIR tasks within the classical music domain. The dataset is composed of over one hour recordings of musical pieces performed by the Colibrì Ensemble at The Spheres recording studio, capturing two canonical works - Tchaikovsky's Romeo and Juliet and Mozart's Symphony No. 40 - along with chromatic scales and solo excerpts for each instrument. The recording setup employed 23 microphones, including close spot, main, and ambient microphones, enabling the creation of realistic stereo mixes with controlled bleeding and providing isolated stems for supervised training of source separation models. In addition, room impulse responses were estimated for each instrument position, offering valuable acoustic characterization of the recording space. We present the dataset structure, acoustic analysis, and baseline evaluations using X-UMX based models for orchestral family separation and microphone debleeding. Results highlight both the potential and the challenges of source separation in complex orchestral scenarios, underscoring the dataset's value for benchmarking and for exploring new approaches to separation, localization, dereverberation, and immersive rendering of classical music.
title The Spheres Dataset: Multitrack Orchestral Recordings for Music Source Separation and Information Retrieval
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2511.21247