Gespeichert in:
| Hauptverfasser: | Han, Wooseok, Kang, Minki, Kim, Changhun, Yang, Eunho |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2412.20155 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping
von: Kang, Minki, et al.
Veröffentlicht: (2023)
von: Kang, Minki, et al.
Veröffentlicht: (2023)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
von: Shin, Seungyoun, et al.
Veröffentlicht: (2025)
von: Shin, Seungyoun, et al.
Veröffentlicht: (2025)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)
Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
von: Lee, Kyowoon, et al.
Veröffentlicht: (2025)
von: Lee, Kyowoon, et al.
Veröffentlicht: (2025)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2025)
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2025)
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
von: Liu, Yutong, et al.
Veröffentlicht: (2025)
von: Liu, Yutong, et al.
Veröffentlicht: (2025)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
von: Xie, Tianxin, et al.
Veröffentlicht: (2025)
von: Xie, Tianxin, et al.
Veröffentlicht: (2025)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
von: Choi, Youngwon, et al.
Veröffentlicht: (2026)
von: Choi, Youngwon, et al.
Veröffentlicht: (2026)
USAT: A Universal Speaker-Adaptive Text-to-Speech Approach
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS
von: Susladkar, Onkar Kishor, et al.
Veröffentlicht: (2024)
von: Susladkar, Onkar Kishor, et al.
Veröffentlicht: (2024)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
von: Deng, Wei, et al.
Veröffentlicht: (2025)
von: Deng, Wei, et al.
Veröffentlicht: (2025)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement
von: Chen, Qianniu, et al.
Veröffentlicht: (2025)
von: Chen, Qianniu, et al.
Veröffentlicht: (2025)
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
von: Chu, Yunji, et al.
Veröffentlicht: (2024)
von: Chu, Yunji, et al.
Veröffentlicht: (2024)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
von: Li, Haitao, et al.
Veröffentlicht: (2026)
von: Li, Haitao, et al.
Veröffentlicht: (2026)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
von: Qi, Xin, et al.
Veröffentlicht: (2024)
von: Qi, Xin, et al.
Veröffentlicht: (2024)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
Adaptive Duration Model for Text Speech Alignment
von: Cao, Junjie
Veröffentlicht: (2025)
von: Cao, Junjie
Veröffentlicht: (2025)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
von: Park, Hyun Joon, et al.
Veröffentlicht: (2024)
von: Park, Hyun Joon, et al.
Veröffentlicht: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
von: Liu, Chenlin, et al.
Veröffentlicht: (2025)
von: Liu, Chenlin, et al.
Veröffentlicht: (2025)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)
von: Shi, Hao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping
von: Kang, Minki, et al.
Veröffentlicht: (2023) -
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
von: Wu, Zhichao, et al.
Veröffentlicht: (2025) -
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025) -
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
von: Shin, Seungyoun, et al.
Veröffentlicht: (2025) -
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)