Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Changhao, Yao, Dongyu, Zhang, Yu, Guo, Wenxiang, Lu, Jingyu, Zhu, Zhiyuan, Zhao, Zhou
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908778589323264
author Pan, Changhao
Yao, Dongyu
Zhang, Yu
Guo, Wenxiang
Lu, Jingyu
Zhu, Zhiyuan
Zhao, Zhou
author_facet Pan, Changhao
Yao, Dongyu
Zhang, Yu
Guo, Wenxiang
Lu, Jingyu
Zhu, Zhiyuan
Zhao, Zhou
contents Recent advances in singing voice synthesis (SVS) have attracted substantial attention from both academia and industry. With the advent of large language models and novel generative paradigms, producing controllable, high-fidelity singing voices has become an attainable goal. Yet the field still lacks a comprehensive survey that systematically analyzes deep-learning-based singing voice synthesis systems and their enabling technologies. To address the aforementioned issue, this survey first categorizes existing systems by task type and then organizes current architectures into two major paradigms: cascaded and end-to-end approaches. Moreover, we provide an in-depth analysis of core technologies, covering singing modeling and control techniques. Finally, we review relevant datasets, annotation tools, and evaluation benchmarks that support training and assessment. In appendix, we introduce training strategies and further discussion of SVS. This survey provides an up-to-date review of the literature on SVS models, which would be a useful reference for both researchers and engineers. Related materials are available at https://github.com/David-Pigeon/SyntheticSingers.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13910
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches
Pan, Changhao
Yao, Dongyu
Zhang, Yu
Guo, Wenxiang
Lu, Jingyu
Zhu, Zhiyuan
Zhao, Zhou
Audio and Speech Processing
Sound
Recent advances in singing voice synthesis (SVS) have attracted substantial attention from both academia and industry. With the advent of large language models and novel generative paradigms, producing controllable, high-fidelity singing voices has become an attainable goal. Yet the field still lacks a comprehensive survey that systematically analyzes deep-learning-based singing voice synthesis systems and their enabling technologies. To address the aforementioned issue, this survey first categorizes existing systems by task type and then organizes current architectures into two major paradigms: cascaded and end-to-end approaches. Moreover, we provide an in-depth analysis of core technologies, covering singing modeling and control techniques. Finally, we review relevant datasets, annotation tools, and evaluation benchmarks that support training and assessment. In appendix, we introduce training strategies and further discussion of SVS. This survey provides an up-to-date review of the literature on SVS models, which would be a useful reference for both researchers and engineers. Related materials are available at https://github.com/David-Pigeon/SyntheticSingers.
title Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2601.13910