MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yutian, Yang, Wanyin, Dai, Zhenrong, Zhang, Yilong, Zhao, Kun, Wang, Hui
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929547893538816
author Wang, Yutian
Yang, Wanyin
Dai, Zhenrong
Zhang, Yilong
Zhao, Kun
Wang, Hui
author_facet Wang, Yutian
Yang, Wanyin
Dai, Zhenrong
Zhang, Yilong
Zhao, Kun
Wang, Hui
contents At present, neural network models show powerful sequence prediction ability and are used in many automatic composition models. In comparison, the way humans compose music is very different from it. Composers usually start by creating musical motifs and then develop them into music through a series of rules. This process ensures that the music has a specific structure and changing pattern. However, it is difficult for neural network models to learn these composition rules from training data, which results in a lack of musicality and diversity in the generated music. This paper posits that integrating the learning capabilities of neural networks with human-derived knowledge may lead to better results. To archive this, we develop the POP909$\_$M dataset, the first to include labels for musical motifs and their variants, providing a basis for mimicking human compositional habits. Building on this, we propose MeloTrans, a text-to-music composition model that employs principles of motif development rules. Our experiments demonstrate that MeloTrans excels beyond existing music generation models and even surpasses Large Language Models (LLMs) like ChatGPT-4. This highlights the importance of merging human insights with neural network capabilities to achieve superior symbolic music generation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13419
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit
Wang, Yutian
Yang, Wanyin
Dai, Zhenrong
Zhang, Yilong
Zhao, Kun
Wang, Hui
Sound
Multimedia
Audio and Speech Processing
At present, neural network models show powerful sequence prediction ability and are used in many automatic composition models. In comparison, the way humans compose music is very different from it. Composers usually start by creating musical motifs and then develop them into music through a series of rules. This process ensures that the music has a specific structure and changing pattern. However, it is difficult for neural network models to learn these composition rules from training data, which results in a lack of musicality and diversity in the generated music. This paper posits that integrating the learning capabilities of neural networks with human-derived knowledge may lead to better results. To archive this, we develop the POP909$\_$M dataset, the first to include labels for musical motifs and their variants, providing a basis for mimicking human compositional habits. Building on this, we propose MeloTrans, a text-to-music composition model that employs principles of motif development rules. Our experiments demonstrate that MeloTrans excels beyond existing music generation models and even surpasses Large Language Models (LLMs) like ChatGPT-4. This highlights the importance of merging human insights with neural network capabilities to achieve superior symbolic music generation.
title MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2410.13419