Transfer the linguistic representations from TTS to accent conversion with non-parallel data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xi, Pei, Jiakun, Xue, Liumeng, Zhang, Mingyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
von: Nguyen, Tuan Nam, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan Nam, et al.
Veröffentlicht: (2024)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
Qwen3-TTS Technical Report
von: Hu, Hangrui, et al.
Veröffentlicht: (2026)
von: Hu, Hangrui, et al.
Veröffentlicht: (2026)
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
von: Chen, Ziqi, et al.
Veröffentlicht: (2025)
von: Chen, Ziqi, et al.
Veröffentlicht: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
von: Tee, Hitomi Jin Ling, et al.
Veröffentlicht: (2025)
von: Tee, Hitomi Jin Ling, et al.
Veröffentlicht: (2025)
MunTTS: A Text-to-Speech System for Mundari
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness
von: Feng, Xincan, et al.
Veröffentlicht: (2024)
von: Feng, Xincan, et al.
Veröffentlicht: (2024)
Multi-interaction TTS toward professional recording reproduction
von: Kanagawa, Hiroki, et al.
Veröffentlicht: (2025)
von: Kanagawa, Hiroki, et al.
Veröffentlicht: (2025)
RWKVTTS: Yet another TTS based on RWKV-7
von: yueyu, Lin, et al.
Veröffentlicht: (2025)
von: yueyu, Lin, et al.
Veröffentlicht: (2025)
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
von: Song, Yaodong, et al.
Veröffentlicht: (2025)
von: Song, Yaodong, et al.
Veröffentlicht: (2025)
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
Word-wise intonation model for cross-language TTS systems
von: A., Tomilov A., et al.
Veröffentlicht: (2024)
von: A., Tomilov A., et al.
Veröffentlicht: (2024)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
von: Roth, Amit, et al.
Veröffentlicht: (2024)
von: Roth, Amit, et al.
Veröffentlicht: (2024)
An investigation of phrase break prediction in an End-to-End TTS system
von: Vadapalli, Anandaswarup
Veröffentlicht: (2023)
von: Vadapalli, Anandaswarup
Veröffentlicht: (2023)
JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
von: Zhou, Fangru, et al.
Veröffentlicht: (2025)
von: Zhou, Fangru, et al.
Veröffentlicht: (2025)
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
von: Liu, Sen, et al.
Veröffentlicht: (2024)
von: Liu, Sen, et al.
Veröffentlicht: (2024)
Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM
von: Yu, Jiawei, et al.
Veröffentlicht: (2024)
von: Yu, Jiawei, et al.
Veröffentlicht: (2024)
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition
von: Chevi, Rendi, et al.
Veröffentlicht: (2024)
von: Chevi, Rendi, et al.
Veröffentlicht: (2024)
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
von: Sanchez, Ariadna, et al.
Veröffentlicht: (2025)
von: Sanchez, Ariadna, et al.
Veröffentlicht: (2025)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
Neighbors and relatives: How do speech embeddings reflect linguistic connections across the world?
von: Törö, Tuukka, et al.
Veröffentlicht: (2025)
von: Törö, Tuukka, et al.
Veröffentlicht: (2025)
CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel Properties
von: Tuttösí, Paige, et al.
Veröffentlicht: (2025)
von: Tuttösí, Paige, et al.
Veröffentlicht: (2025)
A2TTS: TTS for Low Resource Indian Languages
von: Bhadoriya, Ayush Singh, et al.
Veröffentlicht: (2025)
von: Bhadoriya, Ayush Singh, et al.
Veröffentlicht: (2025)
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
von: Lehečka, Jan, et al.
Veröffentlicht: (2024)
von: Lehečka, Jan, et al.
Veröffentlicht: (2024)
Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
von: Fetrat, Mahta, et al.
Veröffentlicht: (2025)
von: Fetrat, Mahta, et al.
Veröffentlicht: (2025)
Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
von: Ueda, Lucas H., et al.
Veröffentlicht: (2024)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
A corpus-based investigation of pitch contours of monosyllabic words in conversational Taiwan Mandarin
von: Jin, Xiaoyun, et al.
Veröffentlicht: (2024)
von: Jin, Xiaoyun, et al.
Veröffentlicht: (2024)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
Semantic enrichment towards efficient speech representations
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
von: Nguyen, Tuan Nam, et al.
Veröffentlicht: (2024) -
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023) -
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023) -
EE-TTS: Emphatic Expressive TTS with Linguistic Information
von: Zhong, Yi, et al.
Veröffentlicht: (2023) -
Qwen3-TTS Technical Report
von: Hu, Hangrui, et al.
Veröffentlicht: (2026)