SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qian, Jiale, Meng, Hao, Zheng, Tian, Zhu, Pengcheng, Lin, Haopeng, Dai, Yuhang, Xie, Hanke, Cao, Wenxiao, Shang, Ruixuan, Wu, Jun, Liu, Hongmei, Wen, Hanlin, Zhao, Jian, Jiang, Zhonglin, Chen, Yong, Yin, Shunshun, Tao, Ming, Wei, Jianguo, Xie, Lei, Wang, Xinsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription
von: Dai, Yuhang, et al.
Veröffentlicht: (2026)
von: Dai, Yuhang, et al.
Veröffentlicht: (2026)
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation
von: Yan, Ruiqi, et al.
Veröffentlicht: (2026)
von: Yan, Ruiqi, et al.
Veröffentlicht: (2026)
SingIt! Singer Voice Transformation
von: Eliav, Amit, et al.
Veröffentlicht: (2024)
von: Eliav, Amit, et al.
Veröffentlicht: (2024)
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
PolySinger: Singing-Voice to Singing-Voice Translation from English to Japanese
von: Antonisen, Silas, et al.
Veröffentlicht: (2024)
von: Antonisen, Silas, et al.
Veröffentlicht: (2024)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
BiSinger: Bilingual Singing Voice Synthesis
von: Zhou, Huali, et al.
Veröffentlicht: (2023)
von: Zhou, Huali, et al.
Veröffentlicht: (2023)
Zero-Shot Duet Singing Voices Separation with Diffusion Models
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself
von: Zeng, Chang, et al.
Veröffentlicht: (2024)
von: Zeng, Chang, et al.
Veröffentlicht: (2024)
StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
YingMusic-Singer-Plus: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
von: Hao, Chunbo, et al.
Veröffentlicht: (2026)
von: Hao, Chunbo, et al.
Veröffentlicht: (2026)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
von: Gu, Ke, et al.
Veröffentlicht: (2025)
von: Gu, Ke, et al.
Veröffentlicht: (2025)
Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches
von: Pan, Changhao, et al.
Veröffentlicht: (2026)
von: Pan, Changhao, et al.
Veröffentlicht: (2026)
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
von: Song, Yulin, et al.
Veröffentlicht: (2024)
von: Song, Yulin, et al.
Veröffentlicht: (2024)
HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2025)
MuSE-SVS: Multi-Singer Emotional Singing Voice Synthesizer that Controls Emotional Intensity
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
von: Zhou, Wangjin, et al.
Veröffentlicht: (2024)
von: Zhou, Wangjin, et al.
Veröffentlicht: (2024)
SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
von: Cui, Jianwei, et al.
Veröffentlicht: (2024)
von: Cui, Jianwei, et al.
Veröffentlicht: (2024)
Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion
von: Sha, Binzhu, et al.
Veröffentlicht: (2023)
von: Sha, Binzhu, et al.
Veröffentlicht: (2023)
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
von: Li, Haiyun, et al.
Veröffentlicht: (2025)
von: Li, Haiyun, et al.
Veröffentlicht: (2025)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
TokSing: Singing Voice Synthesis based on Discrete Tokens
von: Wu, Yuning, et al.
Veröffentlicht: (2024)
von: Wu, Yuning, et al.
Veröffentlicht: (2024)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
von: Kim, Semin, et al.
Veröffentlicht: (2024)
von: Kim, Semin, et al.
Veröffentlicht: (2024)
SingFake: Singing Voice Deepfake Detection
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
S$^2$Voice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion
von: Wang, Ziqian, et al.
Veröffentlicht: (2026)
von: Wang, Ziqian, et al.
Veröffentlicht: (2026)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
von: Li, Junjie, et al.
Veröffentlicht: (2023)
von: Li, Junjie, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription
von: Dai, Yuhang, et al.
Veröffentlicht: (2026) -
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
von: Xie, Hanke, et al.
Veröffentlicht: (2025) -
SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation
von: Yan, Ruiqi, et al.
Veröffentlicht: (2026) -
SingIt! Singer Voice Transformation
von: Eliav, Amit, et al.
Veröffentlicht: (2024) -
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
von: Chen, Shihao, et al.
Veröffentlicht: (2024)