An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios
Fuente:
arXiv
Guardado en:
| Autores principales: | Gong, Cheng, Cooper, Erica, Wang, Xin, Qiang, Chunyu, Geng, Mengzhe, Wells, Dan, Wang, Longbiao, Dang, Jianwu, Tessier, Marc, Pine, Aidan, Richmond, Korin, Yamagishi, Junichi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
por: Gong, Cheng, et al.
Publicado: (2023)
por: Gong, Cheng, et al.
Publicado: (2023)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
Supporting SENCOTEN Language Documentation Efforts with Automatic Speech Recognition
por: Geng, Mengzhe, et al.
Publicado: (2025)
por: Geng, Mengzhe, et al.
Publicado: (2025)
Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and Speaker
por: Gong, Cheng, et al.
Publicado: (2025)
por: Gong, Cheng, et al.
Publicado: (2025)
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
por: Wang, Tianrui, et al.
Publicado: (2025)
por: Wang, Tianrui, et al.
Publicado: (2025)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
por: Zhong, Jinzuomu, et al.
Publicado: (2025)
por: Zhong, Jinzuomu, et al.
Publicado: (2025)
MSR-HuBERT: Self-supervised Pre-training for Adaptation to Multiple Sampling Rates
por: Huang, Zikang, et al.
Publicado: (2026)
por: Huang, Zikang, et al.
Publicado: (2026)
ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
por: Wang, Junyu, et al.
Publicado: (2025)
por: Wang, Junyu, et al.
Publicado: (2025)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
por: Sun, Siqi, et al.
Publicado: (2024)
por: Sun, Siqi, et al.
Publicado: (2024)
Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
por: Tang, Jingjing, et al.
Publicado: (2025)
por: Tang, Jingjing, et al.
Publicado: (2025)
Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches
por: Zeng, Chang, et al.
Publicado: (2024)
por: Zeng, Chang, et al.
Publicado: (2024)
Enriching Multimodal Sentiment Analysis through Textual Emotional Descriptions of Visual-Audio Content
por: Wu, Sheng, et al.
Publicado: (2024)
por: Wu, Sheng, et al.
Publicado: (2024)
Good practices for evaluation of synthesized speech
por: Cooper, Erica, et al.
Publicado: (2025)
por: Cooper, Erica, et al.
Publicado: (2025)
InstructAudio: Unified speech and music generation with natural language instruction
por: Qiang, Chunyu, et al.
Publicado: (2025)
por: Qiang, Chunyu, et al.
Publicado: (2025)
LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
por: Wang, Junyu, et al.
Publicado: (2025)
por: Wang, Junyu, et al.
Publicado: (2025)
Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
por: Wang, Junyu, et al.
Publicado: (2024)
por: Wang, Junyu, et al.
Publicado: (2024)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
por: Chen, Zhengyang, et al.
Publicado: (2024)
por: Chen, Zhengyang, et al.
Publicado: (2024)
AIMDiT: Modality Augmentation and Interaction via Multimodal Dimension Transformation for Emotion Recognition in Conversations
por: Wu, Sheng, et al.
Publicado: (2024)
por: Wu, Sheng, et al.
Publicado: (2024)
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
por: Shu, Yuchun, et al.
Publicado: (2024)
por: Shu, Yuchun, et al.
Publicado: (2024)
SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
por: Qiang, Chunyu, et al.
Publicado: (2025)
por: Qiang, Chunyu, et al.
Publicado: (2025)
Rethinking Discrete Speech Representation Tokens for Accent Generation
por: Zhong, Jinzuomu, et al.
Publicado: (2026)
por: Zhong, Jinzuomu, et al.
Publicado: (2026)
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
por: Qiang, Chunyu, et al.
Publicado: (2024)
por: Qiang, Chunyu, et al.
Publicado: (2024)
CECOR: Correction-oriented synthetic data construction for factual error correction
por: Zhu, Lei, et al.
Publicado: (2026)
por: Zhu, Lei, et al.
Publicado: (2026)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
por: Wang, Tianrui, et al.
Publicado: (2025)
por: Wang, Tianrui, et al.
Publicado: (2025)
Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
por: Zhang, Lin, et al.
Publicado: (2024)
por: Zhang, Lin, et al.
Publicado: (2024)
Rethinking Contrastive Learning in Graph Anomaly Detection: A Clean-View Perspective
por: Jin, Di, et al.
Publicado: (2025)
por: Jin, Di, et al.
Publicado: (2025)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
por: Qiang, Chunyu, et al.
Publicado: (2026)
por: Qiang, Chunyu, et al.
Publicado: (2026)
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
por: Cui, Zhongjian, et al.
Publicado: (2025)
por: Cui, Zhongjian, et al.
Publicado: (2025)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
por: Han, Zhichen, et al.
Publicado: (2024)
por: Han, Zhichen, et al.
Publicado: (2024)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
por: Sanders, Nicholas, et al.
Publicado: (2025)
por: Sanders, Nicholas, et al.
Publicado: (2025)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
por: Sun, Yujia, et al.
Publicado: (2024)
por: Sun, Yujia, et al.
Publicado: (2024)
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
por: Zhong, Jinzuomu, et al.
Publicado: (2024)
por: Zhong, Jinzuomu, et al.
Publicado: (2024)
A Preliminary Case Study on Long-Form In-the-Wild Audio Spoofing Detection
por: Liu, Xuechen, et al.
Publicado: (2024)
por: Liu, Xuechen, et al.
Publicado: (2024)
Towards Data Drift Monitoring for Speech Deepfake Detection in the context of MLOps
por: Wang, Xin, et al.
Publicado: (2025)
por: Wang, Xin, et al.
Publicado: (2025)
Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?
por: Wang, Xin, et al.
Publicado: (2026)
por: Wang, Xin, et al.
Publicado: (2026)
Zero-Day Audio DeepFake Detection via Retrieval Augmentation and Profile Matching
por: Liu, Xuechen, et al.
Publicado: (2025)
por: Liu, Xuechen, et al.
Publicado: (2025)
FakeMark: Deepfake Speech Attribution With Watermarked Artifacts
por: Ge, Wanying, et al.
Publicado: (2025)
por: Ge, Wanying, et al.
Publicado: (2025)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
por: Fu, Ruibo, et al.
Publicado: (2024)
por: Fu, Ruibo, et al.
Publicado: (2024)
Breaking Data Efficiency Dilemma: A Federated and Augmented Learning Framework For Alzheimer's Disease Detection via Speech
por: Wei, Xiao, et al.
Publicado: (2026)
por: Wei, Xiao, et al.
Publicado: (2026)
Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation
por: Luo, Guoqing, et al.
Publicado: (2025)
por: Luo, Guoqing, et al.
Publicado: (2025)
Ejemplares similares
-
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
por: Gong, Cheng, et al.
Publicado: (2023) -
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
por: Wang, Haoyu, et al.
Publicado: (2024) -
Supporting SENCOTEN Language Documentation Efforts with Automatic Speech Recognition
por: Geng, Mengzhe, et al.
Publicado: (2025) -
Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and Speaker
por: Gong, Cheng, et al.
Publicado: (2025) -
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
por: Wang, Tianrui, et al.
Publicado: (2025)