Music Generation using Human-In-The-Loop Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909466458324992 |
|---|---|
| author | Justus, Aju Ani |
| author_facet | Justus, Aju Ani |
| contents | This paper presents an approach that combines Human-In-The-Loop Reinforcement Learning (HITL RL) with principles derived from music theory to facilitate real-time generation of musical compositions. HITL RL, previously employed in diverse applications such as modelling humanoid robot mechanics and enhancing language models, harnesses human feedback to refine the training process. In this study, we develop a HILT RL framework that can leverage the constraints and principles in music theory. In particular, we propose an episodic tabular Q-learning algorithm with an epsilon-greedy exploration policy. The system generates musical tracks (compositions), continuously enhancing its quality through iterative human-in-the-loop feedback. The reward function for this process is the subjective musical taste of the user. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_15304 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Music Generation using Human-In-The-Loop Reinforcement Learning Justus, Aju Ani Sound Artificial Intelligence Human-Computer Interaction Machine Learning Audio and Speech Processing This paper presents an approach that combines Human-In-The-Loop Reinforcement Learning (HITL RL) with principles derived from music theory to facilitate real-time generation of musical compositions. HITL RL, previously employed in diverse applications such as modelling humanoid robot mechanics and enhancing language models, harnesses human feedback to refine the training process. In this study, we develop a HILT RL framework that can leverage the constraints and principles in music theory. In particular, we propose an episodic tabular Q-learning algorithm with an epsilon-greedy exploration policy. The system generates musical tracks (compositions), continuously enhancing its quality through iterative human-in-the-loop feedback. The reward function for this process is the subjective musical taste of the user. |
| title | Music Generation using Human-In-The-Loop Reinforcement Learning |
| topic | Sound Artificial Intelligence Human-Computer Interaction Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2501.15304 |