Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity
Fuente:
arXiv
Salvato in:
| Autori principali: | Ulgen, Ismail Rasim, Hansen, John H. L., Busso, Carlos, Sisman, Berrak |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
Text-to-Speech for Unseen Speakers via Low-Complexity Discrete Unit-Based Frame Selection
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2025)
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2025)
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
di: Zhao, Xiutian, et al.
Pubblicazione: (2026)
di: Zhao, Xiutian, et al.
Pubblicazione: (2026)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
di: Lee, Philip H., et al.
Pubblicazione: (2024)
di: Lee, Philip H., et al.
Pubblicazione: (2024)
ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2026)
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2026)
Can Emotion Fool Anti-spoofing?
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2025)
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2025)
Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
di: Salman, Ali N., et al.
Pubblicazione: (2024)
di: Salman, Ali N., et al.
Pubblicazione: (2024)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2025)
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2025)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2026)
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2026)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
di: Du, Zongyang, et al.
Pubblicazione: (2025)
di: Du, Zongyang, et al.
Pubblicazione: (2025)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based Depression Detection
di: Yeh, Hsiang-Chen, et al.
Pubblicazione: (2026)
di: Yeh, Hsiang-Chen, et al.
Pubblicazione: (2026)
Study on Inter and Intra Speaker Variability in Speaker Recognition
di: Okhotnikov, Anton, et al.
Pubblicazione: (2024)
di: Okhotnikov, Anton, et al.
Pubblicazione: (2024)
Efficient Adapter Tuning of Pre-trained Speech Models for Automatic Speaker Verification
di: Sang, Mufan, et al.
Pubblicazione: (2024)
di: Sang, Mufan, et al.
Pubblicazione: (2024)
emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2024)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2024)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
di: Lin, Chaohao, et al.
Pubblicazione: (2025)
di: Lin, Chaohao, et al.
Pubblicazione: (2025)
Generalizability of Predictive and Generative Speech Enhancement Models to Pathological Speakers
di: Hou, Mingchi, et al.
Pubblicazione: (2025)
di: Hou, Mingchi, et al.
Pubblicazione: (2025)
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
di: Liu, Rui, et al.
Pubblicazione: (2025)
di: Liu, Rui, et al.
Pubblicazione: (2025)
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
di: Zhao, Xiutian, et al.
Pubblicazione: (2026)
di: Zhao, Xiutian, et al.
Pubblicazione: (2026)
Guided Speaker Embedding
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
What Does the Speaker Embedding Encode?
di: Wang, Shuai, et al.
Pubblicazione: (2025)
di: Wang, Shuai, et al.
Pubblicazione: (2025)
UniPET-SPK: A Unified Framework for Parameter-Efficient Tuning of Pre-trained Speech Models for Robust Speaker Verification
di: Sang, Mufan, et al.
Pubblicazione: (2025)
di: Sang, Mufan, et al.
Pubblicazione: (2025)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
di: Lam, Perry, et al.
Pubblicazione: (2022)
di: Lam, Perry, et al.
Pubblicazione: (2022)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
di: Kim, Miseul, et al.
Pubblicazione: (2025)
di: Kim, Miseul, et al.
Pubblicazione: (2025)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
di: Park, Nohil, et al.
Pubblicazione: (2024)
di: Park, Nohil, et al.
Pubblicazione: (2024)
USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
di: Zeng, Bang, et al.
Pubblicazione: (2024)
di: Zeng, Bang, et al.
Pubblicazione: (2024)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
di: Pálka, Petr, et al.
Pubblicazione: (2024)
di: Pálka, Petr, et al.
Pubblicazione: (2024)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
di: Thienpondt, Jenthe, et al.
Pubblicazione: (2024)
di: Thienpondt, Jenthe, et al.
Pubblicazione: (2024)
Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System
di: Li, Ze, et al.
Pubblicazione: (2024)
di: Li, Ze, et al.
Pubblicazione: (2024)
Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech
di: Feng, Fuyuan, et al.
Pubblicazione: (2026)
di: Feng, Fuyuan, et al.
Pubblicazione: (2026)
Versatile audio-visual learning for emotion recognition
di: Goncalves, Lucas, et al.
Pubblicazione: (2023)
di: Goncalves, Lucas, et al.
Pubblicazione: (2023)
Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement
di: Huang, Ziling, et al.
Pubblicazione: (2025)
di: Huang, Ziling, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024) -
Text-to-Speech for Unseen Speakers via Low-Complexity Discrete Unit-Based Frame Selection
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024) -
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
di: Mahapatra, Aurosweta, et al.
Pubblicazione: (2025) -
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
di: Zhao, Xiutian, et al.
Pubblicazione: (2026) -
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
di: Lee, Philip H., et al.
Pubblicazione: (2024)