Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
Fuente:
arXiv
Salvato in:
| Autori principali: | Vecino, Biel Tura, Gabryś, Adam, Mątwicki, Daniel, Pomirski, Andrzej, Iddon, Tom, Cotescu, Marius, Lorenzo-Trueba, Jaime |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Investigating self-supervised features for expressive, multilingual voice conversion
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2025)
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2025)
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2024)
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2024)
End-to-end streaming model for low-latency speech anonymization
di: Quamer, Waris, et al.
Pubblicazione: (2024)
di: Quamer, Waris, et al.
Pubblicazione: (2024)
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
di: Sereda, Taras
Pubblicazione: (2024)
di: Sereda, Taras
Pubblicazione: (2024)
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
di: Tang, Duowei, et al.
Pubblicazione: (2023)
di: Tang, Duowei, et al.
Pubblicazione: (2023)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
di: Huang, Ziling, et al.
Pubblicazione: (2025)
di: Huang, Ziling, et al.
Pubblicazione: (2025)
End-to-end multi-channel speaker extraction and binaural speech synthesis
di: Chi, Cheng, et al.
Pubblicazione: (2024)
di: Chi, Cheng, et al.
Pubblicazione: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
Automated evaluation of children's speech fluency for low-resource languages
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
di: Akkiraju, Bhavana, et al.
Pubblicazione: (2025)
di: Akkiraju, Bhavana, et al.
Pubblicazione: (2025)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
di: Wright, George August, et al.
Pubblicazione: (2023)
di: Wright, George August, et al.
Pubblicazione: (2023)
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
di: San, Nay, et al.
Pubblicazione: (2024)
di: San, Nay, et al.
Pubblicazione: (2024)
Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
di: Vecino, Biel Tura, et al.
Pubblicazione: (2025)
di: Vecino, Biel Tura, et al.
Pubblicazione: (2025)
Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform
di: Kong, Xiangzhu, et al.
Pubblicazione: (2025)
di: Kong, Xiangzhu, et al.
Pubblicazione: (2025)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
di: Cheng, Shanbo, et al.
Pubblicazione: (2025)
di: Cheng, Shanbo, et al.
Pubblicazione: (2025)
Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation
di: Ma, Jianbo, et al.
Pubblicazione: (2026)
di: Ma, Jianbo, et al.
Pubblicazione: (2026)
Enhancement by postfiltering for speech and audio coding in ad-hoc sensor networks
di: Das, Sneha, et al.
Pubblicazione: (2020)
di: Das, Sneha, et al.
Pubblicazione: (2020)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
di: Kesiraju, Santosh, et al.
Pubblicazione: (2023)
di: Kesiraju, Santosh, et al.
Pubblicazione: (2023)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
di: Paraskevopoulos, Georgios, et al.
Pubblicazione: (2024)
di: Paraskevopoulos, Georgios, et al.
Pubblicazione: (2024)
Boosting Diffusion Model for Spectrogram Up-sampling in Text-to-speech: An Empirical Study
di: Zhang, Chong, et al.
Pubblicazione: (2024)
di: Zhang, Chong, et al.
Pubblicazione: (2024)
Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
di: Li, Mohan, et al.
Pubblicazione: (2024)
di: Li, Mohan, et al.
Pubblicazione: (2024)
Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
di: Ren, Bo, et al.
Pubblicazione: (2025)
di: Ren, Bo, et al.
Pubblicazione: (2025)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
di: Raj, Desh
Pubblicazione: (2024)
di: Raj, Desh
Pubblicazione: (2024)
BFA: Real-time Multilingual Text-to-speech Forced Alignment
di: Rehman, Abdul, et al.
Pubblicazione: (2025)
di: Rehman, Abdul, et al.
Pubblicazione: (2025)
Optimizing Byte-level Representation for End-to-end ASR
di: Hsiao, Roger, et al.
Pubblicazione: (2024)
di: Hsiao, Roger, et al.
Pubblicazione: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2024)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
di: Saon, George, et al.
Pubblicazione: (2025)
di: Saon, George, et al.
Pubblicazione: (2025)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
di: Wang, Jianzong, et al.
Pubblicazione: (2023)
di: Wang, Jianzong, et al.
Pubblicazione: (2023)
Deep low-latency joint speech transmission and enhancement over a gaussian channel
di: Bokaei, Mohammad, et al.
Pubblicazione: (2024)
di: Bokaei, Mohammad, et al.
Pubblicazione: (2024)
Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
di: Li, Chia-Yu, et al.
Pubblicazione: (2024)
di: Li, Chia-Yu, et al.
Pubblicazione: (2024)
Boosting keyword spotting through on-device learnable user speech characteristics
di: Cioflan, Cristian, et al.
Pubblicazione: (2024)
di: Cioflan, Cristian, et al.
Pubblicazione: (2024)
Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
di: Mack, Wolfgang, et al.
Pubblicazione: (2025)
di: Mack, Wolfgang, et al.
Pubblicazione: (2025)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
di: Ahmad, Hawraz A., et al.
Pubblicazione: (2024)
di: Ahmad, Hawraz A., et al.
Pubblicazione: (2024)
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
di: Cheng, Shanbo, et al.
Pubblicazione: (2024)
di: Cheng, Shanbo, et al.
Pubblicazione: (2024)
WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion
di: Baoueb, Teysir, et al.
Pubblicazione: (2024)
di: Baoueb, Teysir, et al.
Pubblicazione: (2024)
End-to-end Acoustic-linguistic Emotion and Intent Recognition Enhanced by Semi-supervised Learning
di: Ren, Zhao, et al.
Pubblicazione: (2025)
di: Ren, Zhao, et al.
Pubblicazione: (2025)
Improving endpoint detection in end-to-end streaming ASR for conversational speech
di: C, Anandh, et al.
Pubblicazione: (2025)
di: C, Anandh, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Investigating self-supervised features for expressive, multilingual voice conversion
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2025) -
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2024) -
End-to-end streaming model for low-latency speech anonymization
di: Quamer, Waris, et al.
Pubblicazione: (2024) -
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
di: Sereda, Taras
Pubblicazione: (2024) -
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
di: Tang, Duowei, et al.
Pubblicazione: (2023)