Singer separation for karaoke content generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Hsuan-Yu, Chen, Xuanjun, Jang, Jyh-Shing Roger
Formato: Preprint
Publicado: 2021
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917751504764928
author Lin, Hsuan-Yu
Chen, Xuanjun
Jang, Jyh-Shing Roger
author_facet Lin, Hsuan-Yu
Chen, Xuanjun
Jang, Jyh-Shing Roger
contents Due to the rapid development of deep learning, we can now successfully separate singing voice from mono audio music. However, this separation can only extract human voices from other musical instruments, which is undesirable for karaoke content generation applications that only require the separation of lead singers. For this karaoke application, we need to separate the music containing male and female duets into two vocals, or extract a single lead vocal from the music containing vocal harmony. For this reason, we propose in this article to use a singer separation system, which generates karaoke content for one or two separated lead singers. In particular, we introduced three models for the singer separation task and designed an automatic model selection scheme to distinguish how many lead singers are in the song. We also collected a large enough data set, MIR-SingerSeparation, which has been publicly released to advance the frontier of this research. Our singer separation is most suitable for sentimental ballads and can be directly applied to karaoke content generation. As far as we know, this is the first singer-separation work for real-world karaoke applications.
format Preprint
id arxiv_https___arxiv_org_abs_2110_06707
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Singer separation for karaoke content generation
Lin, Hsuan-Yu
Chen, Xuanjun
Jang, Jyh-Shing Roger
Sound
Multimedia
Audio and Speech Processing
Due to the rapid development of deep learning, we can now successfully separate singing voice from mono audio music. However, this separation can only extract human voices from other musical instruments, which is undesirable for karaoke content generation applications that only require the separation of lead singers. For this karaoke application, we need to separate the music containing male and female duets into two vocals, or extract a single lead vocal from the music containing vocal harmony. For this reason, we propose in this article to use a singer separation system, which generates karaoke content for one or two separated lead singers. In particular, we introduced three models for the singer separation task and designed an automatic model selection scheme to distinguish how many lead singers are in the song. We also collected a large enough data set, MIR-SingerSeparation, which has been publicly released to advance the frontier of this research. Our singer separation is most suitable for sentimental ballads and can be directly applied to karaoke content generation. As far as we know, this is the first singer-separation work for real-world karaoke applications.
title Singer separation for karaoke content generation
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2110.06707