Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Zhengyang, Liu, Xuechen, Cooper, Erica, Yamagishi, Junichi, Qian, Yanmin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
par: Chen, Zhengyang, et autres
Publié: (2024)
par: Chen, Zhengyang, et autres
Publié: (2024)
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
par: Liu, Yun, et autres
Publié: (2024)
par: Liu, Yun, et autres
Publié: (2024)
Target Speaker Extraction with Curriculum Learning
par: Liu, Yun, et autres
Publié: (2024)
par: Liu, Yun, et autres
Publié: (2024)
Explaining Speaker and Spoof Embeddings via Probing
par: Liu, Xuechen, et autres
Publié: (2024)
par: Liu, Xuechen, et autres
Publié: (2024)
From Sharpness to Better Generalization for Speech Deepfake Detection
par: Huang, Wen, et autres
Publié: (2025)
par: Huang, Wen, et autres
Publié: (2025)
Quantifying Source Speaker Leakage in One-to-One Voice Conversion
par: Wellington, Scott, et autres
Publié: (2025)
par: Wellington, Scott, et autres
Publié: (2025)
Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches
par: Zeng, Chang, et autres
Publié: (2024)
par: Zeng, Chang, et autres
Publié: (2024)
Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning
par: Wang, Shuai, et autres
Publié: (2024)
par: Wang, Shuai, et autres
Publié: (2024)
Prototype and Instance Contrastive Learning for Unsupervised Domain Adaptation in Speaker Verification
par: Huang, Wen, et autres
Publié: (2024)
par: Huang, Wen, et autres
Publié: (2024)
Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
par: Chen, Zhengyang, et autres
Publié: (2024)
par: Chen, Zhengyang, et autres
Publié: (2024)
Improving curriculum learning for target speaker extraction with synthetic speakers
par: Liu, Yun, et autres
Publié: (2024)
par: Liu, Yun, et autres
Publié: (2024)
Listen to Extract: Onset-Prompted Target Speaker Extraction
par: Shen, Pengjie, et autres
Publié: (2025)
par: Shen, Pengjie, et autres
Publié: (2025)
Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker Verification
par: Liu, Bei, et autres
Publié: (2024)
par: Liu, Bei, et autres
Publié: (2024)
Mitigating Language Mismatch in SSL-Based Speaker Anonymization
par: Zhang, Zhe, et autres
Publié: (2025)
par: Zhang, Zhe, et autres
Publié: (2025)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
par: Zhu, Xiaoxu, et autres
Publié: (2025)
par: Zhu, Xiaoxu, et autres
Publié: (2025)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
par: Zhang, Leying, et autres
Publié: (2025)
par: Zhang, Leying, et autres
Publié: (2025)
Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
par: Liu, Bei, et autres
Publié: (2024)
par: Liu, Bei, et autres
Publié: (2024)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
par: Park, Nohil, et autres
Publié: (2024)
par: Park, Nohil, et autres
Publié: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
par: Fu, Ruibo, et autres
Publié: (2024)
par: Fu, Ruibo, et autres
Publié: (2024)
LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech
par: Liu, Xuechen, et autres
Publié: (2025)
par: Liu, Xuechen, et autres
Publié: (2025)
Revisiting and Improving Scoring Fusion for Spoofing-aware Speaker Verification Using Compositional Data Analysis
par: Wang, Xin, et autres
Publié: (2024)
par: Wang, Xin, et autres
Publié: (2024)
A Preliminary Case Study on Long-Form In-the-Wild Audio Spoofing Detection
par: Liu, Xuechen, et autres
Publié: (2024)
par: Liu, Xuechen, et autres
Publié: (2024)
SecureSpeech: Prompt-based Speaker and Content Protection
par: Hui, Belinda Soh Hui, et autres
Publié: (2025)
par: Hui, Belinda Soh Hui, et autres
Publié: (2025)
USED: Universal Speaker Extraction and Diarization
par: Ao, Junyi, et autres
Publié: (2023)
par: Ao, Junyi, et autres
Publié: (2023)
Beyond Speaker Identity: Text Guided Target Speech Extraction
par: Huo, Mingyue, et autres
Publié: (2025)
par: Huo, Mingyue, et autres
Publié: (2025)
Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
par: Tang, Jingjing, et autres
Publié: (2025)
par: Tang, Jingjing, et autres
Publié: (2025)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
par: Gong, Cheng, et autres
Publié: (2023)
par: Gong, Cheng, et autres
Publié: (2023)
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
par: Jin, Jiawei, et autres
Publié: (2025)
par: Jin, Jiawei, et autres
Publié: (2025)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
par: Zhou, Xuehao, et autres
Publié: (2024)
par: Zhou, Xuehao, et autres
Publié: (2024)
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
par: Lin, Yuke, et autres
Publié: (2025)
par: Lin, Yuke, et autres
Publié: (2025)
Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
par: Dai, Wang, et autres
Publié: (2025)
par: Dai, Wang, et autres
Publié: (2025)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
par: Horiguchi, Shota, et autres
Publié: (2025)
par: Horiguchi, Shota, et autres
Publié: (2025)
Multi-Level Speaker Representation for Target Speaker Extraction
par: Zhang, Ke, et autres
Publié: (2024)
par: Zhang, Ke, et autres
Publié: (2024)
Efficient Adapter Tuning of Pre-trained Speech Models for Automatic Speaker Verification
par: Sang, Mufan, et autres
Publié: (2024)
par: Sang, Mufan, et autres
Publié: (2024)
Multi-Label Training for Text-Independent Speaker Identification
par: Xue, Yuqi
Publié: (2022)
par: Xue, Yuqi
Publié: (2022)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
par: Kang, Jiawen, et autres
Publié: (2024)
par: Kang, Jiawen, et autres
Publié: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
par: Zhang, Bowen, et autres
Publié: (2025)
par: Zhang, Bowen, et autres
Publié: (2025)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
par: Xie, Hanke, et autres
Publié: (2025)
par: Xie, Hanke, et autres
Publié: (2025)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
par: Jeon, Yejin, et autres
Publié: (2024)
par: Jeon, Yejin, et autres
Publié: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
par: Wang, Weiqing, et autres
Publié: (2025)
par: Wang, Weiqing, et autres
Publié: (2025)
Documents similaires
-
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
par: Chen, Zhengyang, et autres
Publié: (2024) -
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
par: Liu, Yun, et autres
Publié: (2024) -
Target Speaker Extraction with Curriculum Learning
par: Liu, Yun, et autres
Publié: (2024) -
Explaining Speaker and Spoof Embeddings via Probing
par: Liu, Xuechen, et autres
Publié: (2024) -
From Sharpness to Better Generalization for Speech Deepfake Detection
par: Huang, Wen, et autres
Publié: (2025)