kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-Speech
Fuente:
arXiv
Salvato in:
| Autori principali: | Hajal, Karl El, Kulkarni, Ajinkya, Hermann, Enno, -Doss, Mathew Magimai. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
di: Hajal, Karl El, et al.
Pubblicazione: (2025)
di: Hajal, Karl El, et al.
Pubblicazione: (2025)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
di: Hajal, Karl El, et al.
Pubblicazione: (2025)
di: Hajal, Karl El, et al.
Pubblicazione: (2025)
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
di: Sarkar, Eklavya, et al.
Pubblicazione: (2024)
di: Sarkar, Eklavya, et al.
Pubblicazione: (2024)
Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
di: Kulkarni, Ajinkya, et al.
Pubblicazione: (2025)
di: Kulkarni, Ajinkya, et al.
Pubblicazione: (2025)
Private kNN-VC: Interpretable Anonymization of Converted Speech
di: Franzreb, Carlos, et al.
Pubblicazione: (2025)
di: Franzreb, Carlos, et al.
Pubblicazione: (2025)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
di: Elbanna, Gasser, et al.
Pubblicazione: (2024)
di: Elbanna, Gasser, et al.
Pubblicazione: (2024)
Assessment of Personality Dimensions Across Situations Using Conversational Speech
di: Zhang, Alice, et al.
Pubblicazione: (2025)
di: Zhang, Alice, et al.
Pubblicazione: (2025)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
di: Li, Shaojun, et al.
Pubblicazione: (2024)
di: Li, Shaojun, et al.
Pubblicazione: (2024)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
di: Dowerah, Sandipana, et al.
Pubblicazione: (2025)
di: Dowerah, Sandipana, et al.
Pubblicazione: (2025)
kNN For Whisper And Its Effect On Bias And Speaker Adaptation
di: Nachesa, Maya K., et al.
Pubblicazione: (2024)
di: Nachesa, Maya K., et al.
Pubblicazione: (2024)
Feature Representations for Automatic Meerkat Vocalization Classification
di: Mahmoud, Imen Ben, et al.
Pubblicazione: (2024)
di: Mahmoud, Imen Ben, et al.
Pubblicazione: (2024)
kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness Optimization
di: Shao, Keren, et al.
Pubblicazione: (2025)
di: Shao, Keren, et al.
Pubblicazione: (2025)
Zero-Shot Text-to-Speech from Continuous Text Streams
di: Dang, Trung, et al.
Pubblicazione: (2024)
di: Dang, Trung, et al.
Pubblicazione: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
Unveiling Biases while Embracing Sustainability: Assessing the Dual Challenges of Automatic Speech Recognition Systems
di: Kulkarni, Ajinkya, et al.
Pubblicazione: (2025)
di: Kulkarni, Ajinkya, et al.
Pubblicazione: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
di: Lei, Shun, et al.
Pubblicazione: (2023)
di: Lei, Shun, et al.
Pubblicazione: (2023)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
di: Zhang, Leying, et al.
Pubblicazione: (2025)
di: Zhang, Leying, et al.
Pubblicazione: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
di: Sarkar, Eklavya, et al.
Pubblicazione: (2025)
di: Sarkar, Eklavya, et al.
Pubblicazione: (2025)
Text adaptation for speaker verification with speaker-text factorized embeddings
di: Yang, Yexin, et al.
Pubblicazione: (2025)
di: Yang, Yexin, et al.
Pubblicazione: (2025)
Zero-Shot Text-to-Speech for Vietnamese
di: Vu, Thi, et al.
Pubblicazione: (2025)
di: Vu, Thi, et al.
Pubblicazione: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
di: Yang, Mu, et al.
Pubblicazione: (2024)
di: Yang, Mu, et al.
Pubblicazione: (2024)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
di: Zhu, Han, et al.
Pubblicazione: (2025)
di: Zhu, Han, et al.
Pubblicazione: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
di: Chen, Junyang, et al.
Pubblicazione: (2026)
di: Chen, Junyang, et al.
Pubblicazione: (2026)
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech
di: Bak, Taejun, et al.
Pubblicazione: (2024)
di: Bak, Taejun, et al.
Pubblicazione: (2024)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
di: Arora, Akshit, et al.
Pubblicazione: (2024)
di: Arora, Akshit, et al.
Pubblicazione: (2024)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
di: Doan, Khai Duy, et al.
Pubblicazione: (2024)
di: Doan, Khai Duy, et al.
Pubblicazione: (2024)
Zero-Shot Crate Digging: DJ Tool Retrieval Using Speech Activity, Music Structure And CLAP Embeddings
di: Orife, Iroro
Pubblicazione: (2024)
di: Orife, Iroro
Pubblicazione: (2024)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
di: Kunešová, Marie, et al.
Pubblicazione: (2025)
di: Kunešová, Marie, et al.
Pubblicazione: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
di: Nishimura, Yuto, et al.
Pubblicazione: (2024)
di: Nishimura, Yuto, et al.
Pubblicazione: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
di: Li, Xuyuan, et al.
Pubblicazione: (2024)
di: Li, Xuyuan, et al.
Pubblicazione: (2024)
Towards interfacing large language models with ASR systems using confidence measures and prompting
di: Naderi, Maryam, et al.
Pubblicazione: (2024)
di: Naderi, Maryam, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
di: Hajal, Karl El, et al.
Pubblicazione: (2025) -
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
di: Hajal, Karl El, et al.
Pubblicazione: (2025) -
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
di: Sarkar, Eklavya, et al.
Pubblicazione: (2024) -
Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
di: Kulkarni, Ajinkya, et al.
Pubblicazione: (2025) -
Private kNN-VC: Interpretable Anonymization of Converted Speech
di: Franzreb, Carlos, et al.
Pubblicazione: (2025)