An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kunešová, Marie, Hanzlíček, Zdeněk, Matoušek, Jindřich |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
von: Lehečka, Jan, et al.
Veröffentlicht: (2024)
von: Lehečka, Jan, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
von: Pražák, Aleš, et al.
Veröffentlicht: (2025)
von: Pražák, Aleš, et al.
Veröffentlicht: (2025)
MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification
von: Li, Ya, et al.
Veröffentlicht: (2025)
von: Li, Ya, et al.
Veröffentlicht: (2025)
NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
von: Heo, Hyun-Jun, et al.
Veröffentlicht: (2023)
von: Heo, Hyun-Jun, et al.
Veröffentlicht: (2023)
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement
von: Chen, Qianniu, et al.
Veröffentlicht: (2025)
von: Chen, Qianniu, et al.
Veröffentlicht: (2025)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
von: Ueda, Lucas H., et al.
Veröffentlicht: (2024)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2024)
DINO-VITS: Data-Efficient Zero-Shot TTS with Self-Supervised Speaker Verification Loss for Noise Robustness
von: Pankov, Vikentii, et al.
Veröffentlicht: (2023)
von: Pankov, Vikentii, et al.
Veröffentlicht: (2023)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
von: Chen, Zhiyong, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyong, et al.
Veröffentlicht: (2024)
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
von: Meng, Ming, et al.
Veröffentlicht: (2025)
von: Meng, Ming, et al.
Veröffentlicht: (2025)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
Multi-speaker Text-to-speech Training with Speaker Anonymized Data
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
von: Lee, Myungjin, et al.
Veröffentlicht: (2026)
von: Lee, Myungjin, et al.
Veröffentlicht: (2026)
LightCAM: A Fast and Light Implementation of Context-Aware Masking based D-TDNN for Speaker Verification
von: Cao, Di, et al.
Veröffentlicht: (2024)
von: Cao, Di, et al.
Veröffentlicht: (2024)
Intelli-Z: Toward Intelligible Zero-Shot TTS
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
von: HU, Shujie, et al.
Veröffentlicht: (2025)
von: HU, Shujie, et al.
Veröffentlicht: (2025)
GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
Hierarchical speaker representation for target speaker extraction
von: He, Shulin, et al.
Veröffentlicht: (2022)
von: He, Shulin, et al.
Veröffentlicht: (2022)
Learning Emotion-Invariant Speaker Representations for Speaker Verification
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Text adaptation for speaker verification with speaker-text factorized embeddings
von: Yang, Yexin, et al.
Veröffentlicht: (2025)
von: Yang, Yexin, et al.
Veröffentlicht: (2025)
Improving curriculum learning for target speaker extraction with synthetic speakers
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Infant Cry Emotion Recognition Using Improved ECAPA-TDNN with Multiscale Feature Fusion and Attention Enhancement
von: Zhou, Junyu, et al.
Veröffentlicht: (2025)
von: Zhou, Junyu, et al.
Veröffentlicht: (2025)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
von: Lehečka, Jan, et al.
Veröffentlicht: (2024) -
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024) -
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
von: Pražák, Aleš, et al.
Veröffentlicht: (2025) -
MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification
von: Li, Ya, et al.
Veröffentlicht: (2025) -
NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
von: Heo, Hyun-Jun, et al.
Veröffentlicht: (2023)