Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Carbonneau, Marc-André, van Niekerk, Benjamin, Seuté, Hugo, Letendre, Jean-Philippe, Kamper, Herman, Zaïdi, Julian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Spoken-Term Discovery using Discrete Speech Units
di: van Niekerk, Benjamin, et al.
Pubblicazione: (2024)
di: van Niekerk, Benjamin, et al.
Pubblicazione: (2024)
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
di: Kamper, Herman, et al.
Pubblicazione: (2025)
di: Kamper, Herman, et al.
Pubblicazione: (2025)
Revisiting speech segmentation and lexicon learning with better features
di: Kamper, Herman, et al.
Pubblicazione: (2024)
di: Kamper, Herman, et al.
Pubblicazione: (2024)
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
di: Malan, Simon, et al.
Pubblicazione: (2024)
di: Malan, Simon, et al.
Pubblicazione: (2024)
Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
di: Malan, Simon, et al.
Pubblicazione: (2025)
di: Malan, Simon, et al.
Pubblicazione: (2025)
Disentanglement in a GAN for Unconditional Speech Synthesis
di: Baas, Matthew, et al.
Pubblicazione: (2023)
di: Baas, Matthew, et al.
Pubblicazione: (2023)
Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
di: van Rensburg, Kyle Janse, et al.
Pubblicazione: (2026)
di: van Rensburg, Kyle Janse, et al.
Pubblicazione: (2026)
Translating speech with just images
di: Oneata, Dan, et al.
Pubblicazione: (2024)
di: Oneata, Dan, et al.
Pubblicazione: (2024)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
di: Geng, Haopeng, et al.
Pubblicazione: (2024)
di: Geng, Haopeng, et al.
Pubblicazione: (2024)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
di: Avdeeva, Anastasia, et al.
Pubblicazione: (2024)
di: Avdeeva, Anastasia, et al.
Pubblicazione: (2024)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
di: Slabbert, Danel, et al.
Pubblicazione: (2025)
di: Slabbert, Danel, et al.
Pubblicazione: (2025)
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
di: Yin, Chun, et al.
Pubblicazione: (2024)
di: Yin, Chun, et al.
Pubblicazione: (2024)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
di: Zhu, Xinfa, et al.
Pubblicazione: (2023)
di: Zhu, Xinfa, et al.
Pubblicazione: (2023)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
di: Pálka, Petr, et al.
Pubblicazione: (2024)
di: Pálka, Petr, et al.
Pubblicazione: (2024)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
di: Park, Nohil, et al.
Pubblicazione: (2024)
di: Park, Nohil, et al.
Pubblicazione: (2024)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
di: Jacobs, Christiaan, et al.
Pubblicazione: (2025)
di: Jacobs, Christiaan, et al.
Pubblicazione: (2025)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
di: Zhou, Xuehao, et al.
Pubblicazione: (2024)
di: Zhou, Xuehao, et al.
Pubblicazione: (2024)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
Investigation of Speaker Representation for Target-Speaker Speech Processing
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
di: Fu, Ruibo, et al.
Pubblicazione: (2024)
di: Fu, Ruibo, et al.
Pubblicazione: (2024)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
SecureSpeech: Prompt-based Speaker and Content Protection
di: Hui, Belinda Soh Hui, et al.
Pubblicazione: (2025)
di: Hui, Belinda Soh Hui, et al.
Pubblicazione: (2025)
Speaker Anonymisation for Speech-based Suicide Risk Detection
di: Cui, Ziyun, et al.
Pubblicazione: (2025)
di: Cui, Ziyun, et al.
Pubblicazione: (2025)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
di: Kong, Jungil, et al.
Pubblicazione: (2023)
di: Kong, Jungil, et al.
Pubblicazione: (2023)
Speaker-Aware Simulation Improves Conversational Speech Recognition
di: Gedeon, Máté, et al.
Pubblicazione: (2026)
di: Gedeon, Máté, et al.
Pubblicazione: (2026)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping
di: Büthe, Jan, et al.
Pubblicazione: (2023)
di: Büthe, Jan, et al.
Pubblicazione: (2023)
Beyond Speaker Identity: Text Guided Target Speech Extraction
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
di: Zhu, Xiaoxu, et al.
Pubblicazione: (2025)
di: Zhu, Xiaoxu, et al.
Pubblicazione: (2025)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement
di: Huang, Ziling, et al.
Pubblicazione: (2025)
di: Huang, Ziling, et al.
Pubblicazione: (2025)
Phonetic Richness for Improved Automatic Speaker Verification
di: Klein, Nicholas, et al.
Pubblicazione: (2024)
di: Klein, Nicholas, et al.
Pubblicazione: (2024)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2025)
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2025)
Guided Speaker Embedding
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
di: Stahl, Benjamin, et al.
Pubblicazione: (2025)
di: Stahl, Benjamin, et al.
Pubblicazione: (2025)
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
di: Wang, Weiqing, et al.
Pubblicazione: (2024)
di: Wang, Weiqing, et al.
Pubblicazione: (2024)
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
di: Wang, Yuzhu, et al.
Pubblicazione: (2025)
di: Wang, Yuzhu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Spoken-Term Discovery using Discrete Speech Units
di: van Niekerk, Benjamin, et al.
Pubblicazione: (2024) -
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
di: Kamper, Herman, et al.
Pubblicazione: (2025) -
Revisiting speech segmentation and lexicon learning with better features
di: Kamper, Herman, et al.
Pubblicazione: (2024) -
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
di: Malan, Simon, et al.
Pubblicazione: (2024) -
Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
di: Malan, Simon, et al.
Pubblicazione: (2025)