Learning Relationships Between Separate Audio Tracks for Creative Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bujard, Balthazar, Nika, Jérôme, Bevilacqua, Fédéric, Obin, Nicolas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915523023863808
author Bujard, Balthazar
Nika, Jérôme
Bevilacqua, Fédéric
Obin, Nicolas
author_facet Bujard, Balthazar
Nika, Jérôme
Bevilacqua, Fédéric
Obin, Nicolas
contents This paper presents the first step in a research project situated within the field of musical agents. The objective is to achieve, through training, the tuning of the desired musical relationship between a live musical input and a real-time generated musical output, through the curation of a database of separated tracks. We propose an architecture integrating a symbolic decision module capable of learning and exploiting musical relationships from such musical corpus. We detail an offline implementation of this architecture employing Transformers as the decision module, associated with a perception module based on Wav2Vec 2.0, and concatenative synthesis as audio renderer. We present a quantitative evaluation of the decision module's ability to reproduce learned relationships extracted during training. We demonstrate that our decision module can predict a coherent track B when conditioned by its corresponding ''guide'' track A, based on a corpus of paired tracks (A, B).
format Preprint
id arxiv_https___arxiv_org_abs_2509_25296
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Relationships Between Separate Audio Tracks for Creative Applications
Bujard, Balthazar
Nika, Jérôme
Bevilacqua, Fédéric
Obin, Nicolas
Sound
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Audio and Speech Processing
This paper presents the first step in a research project situated within the field of musical agents. The objective is to achieve, through training, the tuning of the desired musical relationship between a live musical input and a real-time generated musical output, through the curation of a database of separated tracks. We propose an architecture integrating a symbolic decision module capable of learning and exploiting musical relationships from such musical corpus. We detail an offline implementation of this architecture employing Transformers as the decision module, associated with a perception module based on Wav2Vec 2.0, and concatenative synthesis as audio renderer. We present a quantitative evaluation of the decision module's ability to reproduce learned relationships extracted during training. We demonstrate that our decision module can predict a coherent track B when conditioned by its corresponding ''guide'' track A, based on a corpus of paired tracks (A, B).
title Learning Relationships Between Separate Audio Tracks for Creative Applications
topic Sound
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2509.25296