Vocal Melody Construction for Persian Lyrics Using LSTM Recurrent Neural Networks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jafari, Farshad, Didehvar, Farzad, Gheibi, Amin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916456667545600
author Jafari, Farshad
Didehvar, Farzad
Gheibi, Amin
author_facet Jafari, Farshad
Didehvar, Farzad
Gheibi, Amin
contents The present paper investigated automatic melody construction for Persian lyrics as an input. It was assumed that there is a phonological correlation between the lyric syllables and the melody in a song. A seq2seq neural network was developed to investigate this assumption, trained on parallel syllable and note sequences in Persian songs to suggest a pleasant melody for a new sequence of syllables. More than 100 pieces of Persian music were collected and converted from the printed version to the digital format due to the lack of a dataset on Persian digital music. Finally, 14 new lyrics were given to the model as input, and the suggested melodies were performed and recorded by music experts to evaluate the trained model. The evaluation was conducted using an audio questionnaire, which more than 170 persons answered. According to the answers about the pleasantness of melody, the system outputs scored an average of 3.005 from 5, while the human-made melodies for the same lyrics obtained an average score of 4.078.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18203
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vocal Melody Construction for Persian Lyrics Using LSTM Recurrent Neural Networks
Jafari, Farshad
Didehvar, Farzad
Gheibi, Amin
Sound
Machine Learning
Audio and Speech Processing
The present paper investigated automatic melody construction for Persian lyrics as an input. It was assumed that there is a phonological correlation between the lyric syllables and the melody in a song. A seq2seq neural network was developed to investigate this assumption, trained on parallel syllable and note sequences in Persian songs to suggest a pleasant melody for a new sequence of syllables. More than 100 pieces of Persian music were collected and converted from the printed version to the digital format due to the lack of a dataset on Persian digital music. Finally, 14 new lyrics were given to the model as input, and the suggested melodies were performed and recorded by music experts to evaluate the trained model. The evaluation was conducted using an audio questionnaire, which more than 170 persons answered. According to the answers about the pleasantness of melody, the system outputs scored an average of 3.005 from 5, while the human-made melodies for the same lyrics obtained an average score of 4.078.
title Vocal Melody Construction for Persian Lyrics Using LSTM Recurrent Neural Networks
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2410.18203