DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | wu, Weihao, Lin, Zhiwei, Zhou, Yixuan, Li, Jingbei, Niu, Rui, Wu, Qinghua, Cao, Songjun, Ma, Long, Wu, Zhiyong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
von: Niu, Rui, et al.
Veröffentlicht: (2025)
von: Niu, Rui, et al.
Veröffentlicht: (2025)
PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
von: Cao, Songjun, et al.
Veröffentlicht: (2025)
von: Cao, Songjun, et al.
Veröffentlicht: (2025)
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
von: Chen, Xueyuan, et al.
Veröffentlicht: (2025)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2025)
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
von: Jin, Jiawei, et al.
Veröffentlicht: (2025)
von: Jin, Jiawei, et al.
Veröffentlicht: (2025)
Generative Expressive Conversational Speech Synthesis
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
Diff-ETS: Learning a Diffusion Probabilistic Model for Electromyography-to-Speech Conversion
von: Ren, Zhao, et al.
Veröffentlicht: (2024)
von: Ren, Zhao, et al.
Veröffentlicht: (2024)
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
von: Li, Yangze, et al.
Veröffentlicht: (2024)
von: Li, Yangze, et al.
Veröffentlicht: (2024)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
von: Zhou, Shuoyi, et al.
Veröffentlicht: (2024)
von: Zhou, Shuoyi, et al.
Veröffentlicht: (2024)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
The THU-HCSI Multi-Speaker Multi-Lingual Few-Shot Voice Cloning System for LIMMITS'24 Challenge
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
Joint Multi-scale Cross-lingual Speaking Style Transfer with Bidirectional Attention Mechanism for Automatic Dubbing
von: Li, Jingbei, et al.
Veröffentlicht: (2023)
von: Li, Jingbei, et al.
Veröffentlicht: (2023)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
von: Ni, Qinke, et al.
Veröffentlicht: (2026)
von: Ni, Qinke, et al.
Veröffentlicht: (2026)
A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
von: Ye, Runchuan, et al.
Veröffentlicht: (2025)
von: Ye, Runchuan, et al.
Veröffentlicht: (2025)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion
von: Sha, Binzhu, et al.
Veröffentlicht: (2023)
von: Sha, Binzhu, et al.
Veröffentlicht: (2023)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks
von: Jin, Xutong, et al.
Veröffentlicht: (2024)
von: Jin, Xutong, et al.
Veröffentlicht: (2024)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
von: Li, Yuke, et al.
Veröffentlicht: (2024)
von: Li, Yuke, et al.
Veröffentlicht: (2024)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis
von: Jia, Zhenqi, et al.
Veröffentlicht: (2024)
von: Jia, Zhenqi, et al.
Veröffentlicht: (2024)
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
von: Li, Fengjin, et al.
Veröffentlicht: (2025)
von: Li, Fengjin, et al.
Veröffentlicht: (2025)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis
von: Gu, Yu, et al.
Veröffentlicht: (2024)
von: Gu, Yu, et al.
Veröffentlicht: (2024)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
von: Niu, Rui, et al.
Veröffentlicht: (2025) -
PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
von: Cao, Songjun, et al.
Veröffentlicht: (2025) -
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
von: Chen, Xueyuan, et al.
Veröffentlicht: (2025) -
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
von: Jin, Jiawei, et al.
Veröffentlicht: (2025) -
Generative Expressive Conversational Speech Synthesis
von: Liu, Rui, et al.
Veröffentlicht: (2024)