Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
Fuente:
arXiv
Guardado en:
| Autores principales: | Borodin, Kirill, Kudryavtsev, Vasiliy, Maslov, Maxim, Vasiliev, Nikita, Gorodnichev, Mikhail, Mkrtchian, Grach |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
por: Borodin, Kirill, et al.
Publicado: (2025)
por: Borodin, Kirill, et al.
Publicado: (2025)
When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus
por: Borodin, Kirill, et al.
Publicado: (2026)
por: Borodin, Kirill, et al.
Publicado: (2026)
Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance
por: Viakhirev, Ivan, et al.
Publicado: (2026)
por: Viakhirev, Ivan, et al.
Publicado: (2026)
Application of ASV for Voice Identification after VC and Duration Predictor Improvement in TTS Models
por: Nikolayevich, Borodin Kirill, et al.
Publicado: (2024)
por: Nikolayevich, Borodin Kirill, et al.
Publicado: (2024)
AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
por: Borodin, Kirill, et al.
Publicado: (2024)
por: Borodin, Kirill, et al.
Publicado: (2024)
Towards Scalable AASIST: Refining Graph Attention for Speech Deepfake Detection
por: Viakhirev, Ivan, et al.
Publicado: (2025)
por: Viakhirev, Ivan, et al.
Publicado: (2025)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
por: Shin, Seungyoun, et al.
Publicado: (2025)
por: Shin, Seungyoun, et al.
Publicado: (2025)
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition
por: Chevi, Rendi, et al.
Publicado: (2024)
por: Chevi, Rendi, et al.
Publicado: (2024)
Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
por: Lee, Kyowoon, et al.
Publicado: (2025)
por: Lee, Kyowoon, et al.
Publicado: (2025)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
por: Han, Wooseok, et al.
Publicado: (2024)
por: Han, Wooseok, et al.
Publicado: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
por: Biadsy, Fadi, et al.
Publicado: (2024)
por: Biadsy, Fadi, et al.
Publicado: (2024)
AutoProsody: A Prosodic Feature Extraction Tool for Indian Languages
por: Thinakaran, Preethi, et al.
Publicado: (2025)
por: Thinakaran, Preethi, et al.
Publicado: (2025)
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
por: Zhong, Jinzuomu, et al.
Publicado: (2023)
por: Zhong, Jinzuomu, et al.
Publicado: (2023)
SponTTS: modeling and transferring spontaneous style for TTS
por: Li, Hanzhao, et al.
Publicado: (2023)
por: Li, Hanzhao, et al.
Publicado: (2023)
E1 TTS: Simple and Fast Non-Autoregressive TTS
por: Liu, Zhijun, et al.
Publicado: (2024)
por: Liu, Zhijun, et al.
Publicado: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
por: Lu, Ye-Xin, et al.
Publicado: (2025)
por: Lu, Ye-Xin, et al.
Publicado: (2025)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
por: Koriyama, Tomoki
Publicado: (2025)
por: Koriyama, Tomoki
Publicado: (2025)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
por: Gudmalwar, Ashishkumar, et al.
Publicado: (2024)
por: Gudmalwar, Ashishkumar, et al.
Publicado: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
por: Eskimez, Sefik Emre, et al.
Publicado: (2024)
por: Eskimez, Sefik Emre, et al.
Publicado: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024)
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
por: Jiang, Yuepeng, et al.
Publicado: (2024)
por: Jiang, Yuepeng, et al.
Publicado: (2024)
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
por: Cui, Wenqian, et al.
Publicado: (2026)
por: Cui, Wenqian, et al.
Publicado: (2026)
DINO-VITS: Data-Efficient Zero-Shot TTS with Self-Supervised Speaker Verification Loss for Noise Robustness
por: Pankov, Vikentii, et al.
Publicado: (2023)
por: Pankov, Vikentii, et al.
Publicado: (2023)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
por: Xue, Heyang, et al.
Publicado: (2025)
por: Xue, Heyang, et al.
Publicado: (2025)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
por: Zhao, Yan, et al.
Publicado: (2024)
por: Zhao, Yan, et al.
Publicado: (2024)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
por: Tang, Haobin, et al.
Publicado: (2024)
por: Tang, Haobin, et al.
Publicado: (2024)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
por: Oh, Hyung-Seok, et al.
Publicado: (2023)
por: Oh, Hyung-Seok, et al.
Publicado: (2023)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
por: Du, Chenpeng, et al.
Publicado: (2021)
por: Du, Chenpeng, et al.
Publicado: (2021)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
por: Kim, Jaehyeon, et al.
Publicado: (2024)
por: Kim, Jaehyeon, et al.
Publicado: (2024)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
por: Guan, Wenhao, et al.
Publicado: (2025)
por: Guan, Wenhao, et al.
Publicado: (2025)
Towards Developing State-of-the-Art TTS Synthesisers for 13 Indian Languages with Signal Processing aided Alignments
por: Prakash, Anusha, et al.
Publicado: (2022)
por: Prakash, Anusha, et al.
Publicado: (2022)
Prosody Analysis of Audiobooks
por: Pethe, Charuta, et al.
Publicado: (2023)
por: Pethe, Charuta, et al.
Publicado: (2023)
Accent-VITS:accent transfer for end-to-end TTS
por: Ma, Linhan, et al.
Publicado: (2023)
por: Ma, Linhan, et al.
Publicado: (2023)
A Dataset for Automatic Assessment of TTS Quality in Spanish
por: Welford, Alejandro Sosa, et al.
Publicado: (2025)
por: Welford, Alejandro Sosa, et al.
Publicado: (2025)
Intelli-Z: Toward Intelligible Zero-Shot TTS
por: Jung, Sunghee, et al.
Publicado: (2024)
por: Jung, Sunghee, et al.
Publicado: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
por: Liu, Huadai, et al.
Publicado: (2023)
por: Liu, Huadai, et al.
Publicado: (2023)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
por: Nguyen, Tan Dat, et al.
Publicado: (2025)
por: Nguyen, Tan Dat, et al.
Publicado: (2025)
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
por: Zeldes, Ella, et al.
Publicado: (2024)
por: Zeldes, Ella, et al.
Publicado: (2024)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
por: He, Xinlu, et al.
Publicado: (2025)
por: He, Xinlu, et al.
Publicado: (2025)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
por: Li, Haoxun, et al.
Publicado: (2025)
por: Li, Haoxun, et al.
Publicado: (2025)
Ejemplares similares
-
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
por: Borodin, Kirill, et al.
Publicado: (2025) -
When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus
por: Borodin, Kirill, et al.
Publicado: (2026) -
Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance
por: Viakhirev, Ivan, et al.
Publicado: (2026) -
Application of ASV for Voice Identification after VC and Duration Predictor Improvement in TTS Models
por: Nikolayevich, Borodin Kirill, et al.
Publicado: (2024) -
AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
por: Borodin, Kirill, et al.
Publicado: (2024)