Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Lehečka, Jan, Hanzlíček, Zdeněk, Matoušek, Jindřich, Tihelka, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
por: Kunešová, Marie, et al.
Publicado: (2025)
por: Kunešová, Marie, et al.
Publicado: (2025)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
por: Jeon, Yejin, et al.
Publicado: (2024)
por: Jeon, Yejin, et al.
Publicado: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
por: Jiang, Ziyue, et al.
Publicado: (2023)
por: Jiang, Ziyue, et al.
Publicado: (2023)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
por: Wu, Zhichao, et al.
Publicado: (2025)
por: Wu, Zhichao, et al.
Publicado: (2025)
Intelli-Z: Toward Intelligible Zero-Shot TTS
por: Jung, Sunghee, et al.
Publicado: (2024)
por: Jung, Sunghee, et al.
Publicado: (2024)
DINO-VITS: Data-Efficient Zero-Shot TTS with Self-Supervised Speaker Verification Loss for Noise Robustness
por: Pankov, Vikentii, et al.
Publicado: (2023)
por: Pankov, Vikentii, et al.
Publicado: (2023)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
por: Kim, Jaehyeon, et al.
Publicado: (2024)
por: Kim, Jaehyeon, et al.
Publicado: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
por: Eskimez, Sefik Emre, et al.
Publicado: (2024)
por: Eskimez, Sefik Emre, et al.
Publicado: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
por: Zhang, Bowen, et al.
Publicado: (2025)
por: Zhang, Bowen, et al.
Publicado: (2025)
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
por: Kunešová, Marie, et al.
Publicado: (2025)
por: Kunešová, Marie, et al.
Publicado: (2025)
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
por: Meng, Ming, et al.
Publicado: (2025)
por: Meng, Ming, et al.
Publicado: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
por: Wang, Chunhui, et al.
Publicado: (2024)
por: Wang, Chunhui, et al.
Publicado: (2024)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
por: Chen, Zhengyang, et al.
Publicado: (2024)
por: Chen, Zhengyang, et al.
Publicado: (2024)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
por: Avdeeva, Anastasia, et al.
Publicado: (2024)
por: Avdeeva, Anastasia, et al.
Publicado: (2024)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
por: Peng, Puyuan, et al.
Publicado: (2025)
por: Peng, Puyuan, et al.
Publicado: (2025)
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
por: Arora, Akshit, et al.
Publicado: (2024)
por: Arora, Akshit, et al.
Publicado: (2024)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
por: Li, Haitao, et al.
Publicado: (2026)
por: Li, Haitao, et al.
Publicado: (2026)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
por: Deng, Wei, et al.
Publicado: (2025)
por: Deng, Wei, et al.
Publicado: (2025)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
por: Lee, Seokgi, et al.
Publicado: (2025)
por: Lee, Seokgi, et al.
Publicado: (2025)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
por: Li, Yinghao Aaron, et al.
Publicado: (2024)
por: Li, Yinghao Aaron, et al.
Publicado: (2024)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
The THU-HCSI Multi-Speaker Multi-Lingual Few-Shot Voice Cloning System for LIMMITS'24 Challenge
por: Zhou, Yixuan, et al.
Publicado: (2024)
por: Zhou, Yixuan, et al.
Publicado: (2024)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
por: Zhu, Xiaoxu, et al.
Publicado: (2025)
por: Zhu, Xiaoxu, et al.
Publicado: (2025)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
por: Fu, Ruibo, et al.
Publicado: (2024)
por: Fu, Ruibo, et al.
Publicado: (2024)
Zero-Shot Multi-Lingual Speaker Verification in Clinical Trials
por: Akram, Ali, et al.
Publicado: (2024)
por: Akram, Ali, et al.
Publicado: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
por: Nespoli, Francesco, et al.
Publicado: (2024)
por: Nespoli, Francesco, et al.
Publicado: (2024)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
por: Yang, Mu, et al.
Publicado: (2024)
por: Yang, Mu, et al.
Publicado: (2024)
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
por: Jiang, Ziyue, et al.
Publicado: (2025)
por: Jiang, Ziyue, et al.
Publicado: (2025)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
por: Kim, Taesoo, et al.
Publicado: (2025)
por: Kim, Taesoo, et al.
Publicado: (2025)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
por: Chen, Zhengyang, et al.
Publicado: (2024)
por: Chen, Zhengyang, et al.
Publicado: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
por: Li, Xuyuan, et al.
Publicado: (2024)
por: Li, Xuyuan, et al.
Publicado: (2024)
Zero-Shot Text-to-Speech from Continuous Text Streams
por: Dang, Trung, et al.
Publicado: (2024)
por: Dang, Trung, et al.
Publicado: (2024)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
por: Alsayegh, Ali, et al.
Publicado: (2025)
por: Alsayegh, Ali, et al.
Publicado: (2025)
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
por: Choi, Youngwon, et al.
Publicado: (2026)
por: Choi, Youngwon, et al.
Publicado: (2026)
Pre-Finetuning for Few-Shot Emotional Speech Recognition
por: Chen, Maximillian, et al.
Publicado: (2023)
por: Chen, Maximillian, et al.
Publicado: (2023)
The Universal Personalizer: Few-Shot Dysarthric Speech Recognition via Meta-Learning
por: Agarwal, Dhruuv, et al.
Publicado: (2025)
por: Agarwal, Dhruuv, et al.
Publicado: (2025)
Voice Impression Control in Zero-Shot TTS
por: Fujita, Kenichi, et al.
Publicado: (2025)
por: Fujita, Kenichi, et al.
Publicado: (2025)
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
por: Zhou, Shuoyi, et al.
Publicado: (2024)
por: Zhou, Shuoyi, et al.
Publicado: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
por: Jiang, Yuepeng, et al.
Publicado: (2024)
por: Jiang, Yuepeng, et al.
Publicado: (2024)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
por: Zhu, Han, et al.
Publicado: (2025)
por: Zhu, Han, et al.
Publicado: (2025)
Ejemplares similares
-
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
por: Kunešová, Marie, et al.
Publicado: (2025) -
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
por: Jeon, Yejin, et al.
Publicado: (2024) -
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
por: Jiang, Ziyue, et al.
Publicado: (2023) -
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
por: Wu, Zhichao, et al.
Publicado: (2025) -
Intelli-Z: Toward Intelligible Zero-Shot TTS
por: Jung, Sunghee, et al.
Publicado: (2024)