Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
Fuente:
arXiv
Salvato in:
| Autori principali: | Răgman, Teodora, Stan, Adriana |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
di: Răgman, Teodora, et al.
Pubblicazione: (2026)
di: Răgman, Teodora, et al.
Pubblicazione: (2026)
Arabic TTS with FastPitch: Reproducible Baselines, Adversarial Training, and Oversmoothing Analysis
di: Nippert, Lars
Pubblicazione: (2025)
di: Nippert, Lars
Pubblicazione: (2025)
Incremental FastPitch: Chunk-based High Quality Text to Speech
di: Du, Muyang, et al.
Pubblicazione: (2024)
di: Du, Muyang, et al.
Pubblicazione: (2024)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
di: Huang, Ziling, et al.
Pubblicazione: (2025)
di: Huang, Ziling, et al.
Pubblicazione: (2025)
End-to-end multi-channel speaker extraction and binaural speech synthesis
di: Chi, Cheng, et al.
Pubblicazione: (2024)
di: Chi, Cheng, et al.
Pubblicazione: (2024)
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
di: Ogun, Sewade, et al.
Pubblicazione: (2024)
di: Ogun, Sewade, et al.
Pubblicazione: (2024)
Text adaptation for speaker verification with speaker-text factorized embeddings
di: Yang, Yexin, et al.
Pubblicazione: (2025)
di: Yang, Yexin, et al.
Pubblicazione: (2025)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
di: Tang, Duowei, et al.
Pubblicazione: (2023)
di: Tang, Duowei, et al.
Pubblicazione: (2023)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
di: Zhang, Yiru, et al.
Pubblicazione: (2025)
di: Zhang, Yiru, et al.
Pubblicazione: (2025)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
di: Han, Runduo, et al.
Pubblicazione: (2024)
di: Han, Runduo, et al.
Pubblicazione: (2024)
Improving fairness in speaker verification via Group-adapted Fusion Network
di: Shen, Hua, et al.
Pubblicazione: (2022)
di: Shen, Hua, et al.
Pubblicazione: (2022)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
di: An, Keyu, et al.
Pubblicazione: (2025)
di: An, Keyu, et al.
Pubblicazione: (2025)
Online speaker diarization of meetings guided by speech separation
di: Gruttadauria, Elio, et al.
Pubblicazione: (2024)
di: Gruttadauria, Elio, et al.
Pubblicazione: (2024)
Multi-speaker Text-to-speech Training with Speaker Anonymized Data
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
Hierarchical speaker representation for target speaker extraction
di: He, Shulin, et al.
Pubblicazione: (2022)
di: He, Shulin, et al.
Pubblicazione: (2022)
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
di: Terashima, Ryo, et al.
Pubblicazione: (2025)
di: Terashima, Ryo, et al.
Pubblicazione: (2025)
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
di: Pîrlogeanu, Gabriel, et al.
Pubblicazione: (2026)
di: Pîrlogeanu, Gabriel, et al.
Pubblicazione: (2026)
Teaching the Teachers: Boosting unsupervised domain adaptation in speech recognition by ensemble update
di: Ahmad, Rehan, et al.
Pubblicazione: (2026)
di: Ahmad, Rehan, et al.
Pubblicazione: (2026)
Simi-SFX: A similarity-based conditioning method for controllable sound effect synthesis
di: Liu, Yunyi, et al.
Pubblicazione: (2024)
di: Liu, Yunyi, et al.
Pubblicazione: (2024)
Triage knowledge distillation for speaker verification
di: Kim, Ju-ho, et al.
Pubblicazione: (2026)
di: Kim, Ju-ho, et al.
Pubblicazione: (2026)
Privacy-oriented manipulation of speaker representations
di: Teixeira, Francisco, et al.
Pubblicazione: (2023)
di: Teixeira, Francisco, et al.
Pubblicazione: (2023)
Improving curriculum learning for target speaker extraction with synthetic speakers
di: Liu, Yun, et al.
Pubblicazione: (2024)
di: Liu, Yun, et al.
Pubblicazione: (2024)
Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
di: Li, Xuyuan, et al.
Pubblicazione: (2023)
di: Li, Xuyuan, et al.
Pubblicazione: (2023)
Challenging margin-based speaker embedding extractors by using the variational information bottleneck
di: Stafylakis, Themos, et al.
Pubblicazione: (2024)
di: Stafylakis, Themos, et al.
Pubblicazione: (2024)
Tandem spoofing-robust automatic speaker verification based on time-domain embeddings
di: Weizman, Avishai, et al.
Pubblicazione: (2024)
di: Weizman, Avishai, et al.
Pubblicazione: (2024)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
di: Gubian, Michele, et al.
Pubblicazione: (2025)
di: Gubian, Michele, et al.
Pubblicazione: (2025)
SCDiar: a streaming diarization system based on speaker change detection and speech recognition
di: Zheng, Naijun, et al.
Pubblicazione: (2025)
di: Zheng, Naijun, et al.
Pubblicazione: (2025)
Curriculum learning for self-supervised speaker verification
di: Heo, Hee-Soo, et al.
Pubblicazione: (2022)
di: Heo, Hee-Soo, et al.
Pubblicazione: (2022)
Quantifying the effect of speech pathology on automatic and human speaker verification
di: Halpern, Bence Mark, et al.
Pubblicazione: (2024)
di: Halpern, Bence Mark, et al.
Pubblicazione: (2024)
An adaptive filter bank based neural network approach for time delay estimation and speech enhancement
di: Ma, Lu
Pubblicazione: (2025)
di: Ma, Lu
Pubblicazione: (2025)
WavLM model ensemble for audio deepfake detection
di: Combei, David, et al.
Pubblicazione: (2024)
di: Combei, David, et al.
Pubblicazione: (2024)
TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes
di: Stan, Adriana, et al.
Pubblicazione: (2025)
di: Stan, Adriana, et al.
Pubblicazione: (2025)
AlignNet: Learning dataset score alignment functions to enable better training of speech quality estimators
di: Pieper, Jaden, et al.
Pubblicazione: (2024)
di: Pieper, Jaden, et al.
Pubblicazione: (2024)
TBDM-Net: Bidirectional Dense Networks with Gender Information for Speech Emotion Recognition
di: Striletchi, Vlad, et al.
Pubblicazione: (2024)
di: Striletchi, Vlad, et al.
Pubblicazione: (2024)
Investigation of perception inconsistency in speaker embedding for asynchronous voice anonymization
di: Wang, Rui, et al.
Pubblicazione: (2025)
di: Wang, Rui, et al.
Pubblicazione: (2025)
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
di: Rautenberg, Frederik, et al.
Pubblicazione: (2026)
di: Rautenberg, Frederik, et al.
Pubblicazione: (2026)
Why disentanglement-based speaker anonymization systems fail at preserving emotions?
di: Gaznepoglu, Ünal Ege, et al.
Pubblicazione: (2025)
di: Gaznepoglu, Ünal Ege, et al.
Pubblicazione: (2025)
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024)
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024)
EasyEyes: Online hearing research using speakers calibrated by phones
di: Vican, Ivan, et al.
Pubblicazione: (2025)
di: Vican, Ivan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
di: Răgman, Teodora, et al.
Pubblicazione: (2026) -
Arabic TTS with FastPitch: Reproducible Baselines, Adversarial Training, and Oversmoothing Analysis
di: Nippert, Lars
Pubblicazione: (2025) -
Incremental FastPitch: Chunk-based High Quality Text to Speech
di: Du, Muyang, et al.
Pubblicazione: (2024) -
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
di: Huang, Ziling, et al.
Pubblicazione: (2025) -
End-to-end multi-channel speaker extraction and binaural speech synthesis
di: Chi, Cheng, et al.
Pubblicazione: (2024)