PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gu, Ke, Wu, Zhicong, Bai, Peng, Qiao, Sitong, Jiang, Zhiqi, Lu, Junchen, Shi, Xiaodong, Qian, Xinyuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908562005950464
author Gu, Ke
Wu, Zhicong
Bai, Peng
Qiao, Sitong
Jiang, Zhiqi
Lu, Junchen
Shi, Xiaodong
Qian, Xinyuan
author_facet Gu, Ke
Wu, Zhicong
Bai, Peng
Qiao, Sitong
Jiang, Zhiqi
Lu, Junchen
Shi, Xiaodong
Qian, Xinyuan
contents Existing singing voice synthesis (SVS) models largely rely on fine-grained, phoneme-level durations, which limits their practical application. These methods overlook the complementary role of visual information in duration prediction.To address these issues, we propose PerformSinger, a pioneering multimodal SVS framework, which incorporates lip cues from video as a visual modality, enabling high-quality "duration-free" singing voice synthesis. PerformSinger comprises parallel multi-branch multimodal encoders, a feature fusion module, a duration and variational prediction network, a mel-spectrogram decoder and a vocoder. The fusion module, composed of adapter and fusion blocks, employs a progressive fusion strategy within an aligned semantic space to produce high-quality multimodal feature representations, thereby enabling accurate duration prediction and high-fidelity audio synthesis. To facilitate the research, we design, collect and annotate a novel SVS dataset involving synchronized video streams and precise phoneme-level manual annotations. Extensive experiments demonstrate the state-of-the-art performance of our proposal in both subjective and objective evaluations. The code and dataset will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22718
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
Gu, Ke
Wu, Zhicong
Bai, Peng
Qiao, Sitong
Jiang, Zhiqi
Lu, Junchen
Shi, Xiaodong
Qian, Xinyuan
Audio and Speech Processing
Multimedia
Sound
Existing singing voice synthesis (SVS) models largely rely on fine-grained, phoneme-level durations, which limits their practical application. These methods overlook the complementary role of visual information in duration prediction.To address these issues, we propose PerformSinger, a pioneering multimodal SVS framework, which incorporates lip cues from video as a visual modality, enabling high-quality "duration-free" singing voice synthesis. PerformSinger comprises parallel multi-branch multimodal encoders, a feature fusion module, a duration and variational prediction network, a mel-spectrogram decoder and a vocoder. The fusion module, composed of adapter and fusion blocks, employs a progressive fusion strategy within an aligned semantic space to produce high-quality multimodal feature representations, thereby enabling accurate duration prediction and high-fidelity audio synthesis. To facilitate the research, we design, collect and annotate a novel SVS dataset involving synchronized video streams and precise phoneme-level manual annotations. Extensive experiments demonstrate the state-of-the-art performance of our proposal in both subjective and objective evaluations. The code and dataset will be publicly available.
title PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
topic Audio and Speech Processing
Multimedia
Sound
url https://arxiv.org/abs/2509.22718