Marco-Voice Technical Report

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Fengping, Lyu, Chenyang, Ni, Xuanfan, Sun, Haoqin, Li, Qingjuan, Qian, Zhiqiang, Li, Haijun, Wang, Longyue, Xu, Zhao, Luo, Weihua, Zhang, Kaifu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913990454542336
author Tian, Fengping
Lyu, Chenyang
Ni, Xuanfan
Sun, Haoqin
Li, Qingjuan
Qian, Zhiqiang
Li, Haijun
Wang, Longyue
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
author_facet Tian, Fengping
Lyu, Chenyang
Ni, Xuanfan
Sun, Haoqin
Li, Qingjuan
Qian, Zhiqiang
Li, Haijun
Wang, Longyue
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
contents This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in achieving highly expressive, controllable, and natural speech generation that faithfully preserves speaker identity across diverse linguistic and emotional contexts. Our approach introduces an effective speaker-emotion disentanglement mechanism with in-batch contrastive learning, enabling independent manipulation of speaker identity and eemotional style, as well as rotational emotional embedding integration method for smooth emotion control. To support comprehensive training and evaluation, we construct CSEMOTIONS, a high-quality emotional speech dataset containing 10 hours of Mandarin speech from six professional speakers across seven emotional categories. Extensive experiments demonstrate that our system, Marco-Voice, achieves substantial improvements in both objective and subjective metrics. Comprehensive evaluations and analysis were conducted, results show that MarcoVoice delivers competitive performance in terms of speech clarity and emotional richness, representing a substantial advance in the field of expressive neural speech synthesis. Our code and dataset are publicly available at https://github.com/AIDC-AI/Marco-Voice and https://huggingface.co/datasets/AIDC-AI/CSEMOTIONS respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Marco-Voice Technical Report
Tian, Fengping
Lyu, Chenyang
Ni, Xuanfan
Sun, Haoqin
Li, Qingjuan
Qian, Zhiqiang
Li, Haijun
Wang, Longyue
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Computation and Language
Sound
Audio and Speech Processing
This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in achieving highly expressive, controllable, and natural speech generation that faithfully preserves speaker identity across diverse linguistic and emotional contexts. Our approach introduces an effective speaker-emotion disentanglement mechanism with in-batch contrastive learning, enabling independent manipulation of speaker identity and eemotional style, as well as rotational emotional embedding integration method for smooth emotion control. To support comprehensive training and evaluation, we construct CSEMOTIONS, a high-quality emotional speech dataset containing 10 hours of Mandarin speech from six professional speakers across seven emotional categories. Extensive experiments demonstrate that our system, Marco-Voice, achieves substantial improvements in both objective and subjective metrics. Comprehensive evaluations and analysis were conducted, results show that MarcoVoice delivers competitive performance in terms of speech clarity and emotional richness, representing a substantial advance in the field of expressive neural speech synthesis. Our code and dataset are publicly available at https://github.com/AIDC-AI/Marco-Voice and https://huggingface.co/datasets/AIDC-AI/CSEMOTIONS respectively.
title Marco-Voice Technical Report
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.02038