Yin-Yang: Developing Motifs With Long-Term Structure And Controllability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhandari, Keshav, Wiggins, Geraint A., Colton, Simon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915128048353280
author Bhandari, Keshav
Wiggins, Geraint A.
Colton, Simon
author_facet Bhandari, Keshav
Wiggins, Geraint A.
Colton, Simon
contents Transformer models have made great strides in generating symbolically represented music with local coherence. However, controlling the development of motifs in a structured way with global form remains an open research area. One of the reasons for this challenge is due to the note-by-note autoregressive generation of such models, which lack the ability to correct themselves after deviations from the motif. In addition, their structural performance on datasets with shorter durations has not been studied in the literature. In this study, we propose Yin-Yang, a framework consisting of a phrase generator, phrase refiner, and phrase selector models for the development of motifs into melodies with long-term structure and controllability. The phrase refiner is trained on a novel corruption-refinement strategy which allows it to produce melodic and rhythmic variations of an original motif at generation time, thereby rectifying deviations of the phrase generator. We also introduce a new objective evaluation metric for quantifying how smoothly the motif manifests itself within the piece. Evaluation results show that our model achieves better performance compared to state-of-the-art transformer models while having the advantage of being controllable and making the generated musical structure semi-interpretable, paving the way for musical analysis. Our code and demo page can be found at https://github.com/keshavbhandari/yinyang.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17759
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Yin-Yang: Developing Motifs With Long-Term Structure And Controllability
Bhandari, Keshav
Wiggins, Geraint A.
Colton, Simon
Sound
Artificial Intelligence
Symbolic Computation
Transformer models have made great strides in generating symbolically represented music with local coherence. However, controlling the development of motifs in a structured way with global form remains an open research area. One of the reasons for this challenge is due to the note-by-note autoregressive generation of such models, which lack the ability to correct themselves after deviations from the motif. In addition, their structural performance on datasets with shorter durations has not been studied in the literature. In this study, we propose Yin-Yang, a framework consisting of a phrase generator, phrase refiner, and phrase selector models for the development of motifs into melodies with long-term structure and controllability. The phrase refiner is trained on a novel corruption-refinement strategy which allows it to produce melodic and rhythmic variations of an original motif at generation time, thereby rectifying deviations of the phrase generator. We also introduce a new objective evaluation metric for quantifying how smoothly the motif manifests itself within the piece. Evaluation results show that our model achieves better performance compared to state-of-the-art transformer models while having the advantage of being controllable and making the generated musical structure semi-interpretable, paving the way for musical analysis. Our code and demo page can be found at https://github.com/keshavbhandari/yinyang.
title Yin-Yang: Developing Motifs With Long-Term Structure And Controllability
topic Sound
Artificial Intelligence
Symbolic Computation
url https://arxiv.org/abs/2501.17759