Music Generation using Human-In-The-Loop Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Justus, Aju Ani
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909466458324992
author Justus, Aju Ani
author_facet Justus, Aju Ani
contents This paper presents an approach that combines Human-In-The-Loop Reinforcement Learning (HITL RL) with principles derived from music theory to facilitate real-time generation of musical compositions. HITL RL, previously employed in diverse applications such as modelling humanoid robot mechanics and enhancing language models, harnesses human feedback to refine the training process. In this study, we develop a HILT RL framework that can leverage the constraints and principles in music theory. In particular, we propose an episodic tabular Q-learning algorithm with an epsilon-greedy exploration policy. The system generates musical tracks (compositions), continuously enhancing its quality through iterative human-in-the-loop feedback. The reward function for this process is the subjective musical taste of the user.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Music Generation using Human-In-The-Loop Reinforcement Learning
Justus, Aju Ani
Sound
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Audio and Speech Processing
This paper presents an approach that combines Human-In-The-Loop Reinforcement Learning (HITL RL) with principles derived from music theory to facilitate real-time generation of musical compositions. HITL RL, previously employed in diverse applications such as modelling humanoid robot mechanics and enhancing language models, harnesses human feedback to refine the training process. In this study, we develop a HILT RL framework that can leverage the constraints and principles in music theory. In particular, we propose an episodic tabular Q-learning algorithm with an epsilon-greedy exploration policy. The system generates musical tracks (compositions), continuously enhancing its quality through iterative human-in-the-loop feedback. The reward function for this process is the subjective musical taste of the user.
title Music Generation using Human-In-The-Loop Reinforcement Learning
topic Sound
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2501.15304