ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912263167803392 |
|---|---|
| author | Song, Yulin Sang, Guorui Yu, Jing Xiao, Chuangbai |
| author_facet | Song, Yulin Sang, Guorui Yu, Jing Xiao, Chuangbai |
| contents | Singing voice synthesis (SVS) system is expected to generate high-fidelity singing voice from given music scores (lyrics, duration and pitch). Recently, diffusion models have performed well in this field. However, sacrificing inference speed to exchange with high-quality sample generation limits its application scenarios. In order to obtain high quality synthetic singing voice more efficiently, we propose a singing voice synthesis method based on the consistency model, ConSinger, to achieve high-fidelity singing voice synthesis with minimal steps. The model is trained by applying consistency constraint and the generation quality is greatly improved at the expense of a small amount of inference speed. Our experiments show that ConSinger is highly competitive with the baseline model in terms of generation speed and quality. Audio samples are available at https://keylxiao.github.io/consinger. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_15342 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps Song, Yulin Sang, Guorui Yu, Jing Xiao, Chuangbai Sound Machine Learning Audio and Speech Processing Singing voice synthesis (SVS) system is expected to generate high-fidelity singing voice from given music scores (lyrics, duration and pitch). Recently, diffusion models have performed well in this field. However, sacrificing inference speed to exchange with high-quality sample generation limits its application scenarios. In order to obtain high quality synthetic singing voice more efficiently, we propose a singing voice synthesis method based on the consistency model, ConSinger, to achieve high-fidelity singing voice synthesis with minimal steps. The model is trained by applying consistency constraint and the generation quality is greatly improved at the expense of a small amount of inference speed. Our experiments show that ConSinger is highly competitive with the baseline model in terms of generation speed and quality. Audio samples are available at https://keylxiao.github.io/consinger. |
| title | ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps |
| topic | Sound Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2410.15342 |