kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hajal, Karl El, Kulkarni, Ajinkya, Hermann, Enno, -Doss, Mathew Magimai. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
von: Hajal, Karl El, et al.
Veröffentlicht: (2025)
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2024)
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2024)
Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2025)
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2025)
Private kNN-VC: Interpretable Anonymization of Converted Speech
von: Franzreb, Carlos, et al.
Veröffentlicht: (2025)
von: Franzreb, Carlos, et al.
Veröffentlicht: (2025)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024)
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024)
Assessment of Personality Dimensions Across Situations Using Conversational Speech
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
von: Dowerah, Sandipana, et al.
Veröffentlicht: (2025)
von: Dowerah, Sandipana, et al.
Veröffentlicht: (2025)
kNN For Whisper And Its Effect On Bias And Speaker Adaptation
von: Nachesa, Maya K., et al.
Veröffentlicht: (2024)
von: Nachesa, Maya K., et al.
Veröffentlicht: (2024)
Feature Representations for Automatic Meerkat Vocalization Classification
von: Mahmoud, Imen Ben, et al.
Veröffentlicht: (2024)
von: Mahmoud, Imen Ben, et al.
Veröffentlicht: (2024)
kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization
von: Shao, Keren, et al.
Veröffentlicht: (2025)
von: Shao, Keren, et al.
Veröffentlicht: (2025)
Zero-Shot Text-to-Speech from Continuous Text Streams
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
Unveiling Biases while Embracing Sustainability: Assessing the Dual Challenges of Automatic Speech Recognition Systems
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2025)
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2025)
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2025)
Text adaptation for speaker verification with speaker-text factorized embeddings
von: Yang, Yexin, et al.
Veröffentlicht: (2025)
von: Yang, Yexin, et al.
Veröffentlicht: (2025)
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)
von: Vu, Thi, et al.
Veröffentlicht: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
von: Yang, Mu, et al.
Veröffentlicht: (2024)
von: Yang, Mu, et al.
Veröffentlicht: (2024)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech
von: Bak, Taejun, et al.
Veröffentlicht: (2024)
von: Bak, Taejun, et al.
Veröffentlicht: (2024)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
Zero-Shot Crate Digging: DJ Tool Retrieval Using Speech Activity, Music Structure And CLAP Embeddings
von: Orife, Iroro
Veröffentlicht: (2024)
von: Orife, Iroro
Veröffentlicht: (2024)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
von: Nishimura, Yuto, et al.
Veröffentlicht: (2024)
von: Nishimura, Yuto, et al.
Veröffentlicht: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
Towards interfacing large language models with ASR systems using confidence measures and prompting
von: Naderi, Maryam, et al.
Veröffentlicht: (2024)
von: Naderi, Maryam, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
von: Hajal, Karl El, et al.
Veröffentlicht: (2025) -
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
von: Hajal, Karl El, et al.
Veröffentlicht: (2025) -
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2024) -
Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2025) -
Private kNN-VC: Interpretable Anonymization of Converted Speech
von: Franzreb, Carlos, et al.
Veröffentlicht: (2025)