You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel Properties
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tuttösí, Paige, Yeung, H. Henny, Wang, Yue, Aucouturier, Jean-Julien, Lim, Angelica |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mmm whatcha say? Uncovering distal and proximal context effects in first and second-language word perception using psychophysical reverse correlation
par: Tuttösí, Paige, et autres
Publié: (2024)
par: Tuttösí, Paige, et autres
Publié: (2024)
I Know You're Listening: Adaptive Voice for HRI
par: Tuttösí, Paige
Publié: (2025)
par: Tuttösí, Paige
Publié: (2025)
Quantification of Tenseness in English and Japanese Tense-Lax Vowels: A Lagrangian Model with Indicator θ1 and Force of Tenseness Ftense(t)
par: Ishizaki, Tatsuya
Publié: (2025)
par: Ishizaki, Tatsuya
Publié: (2025)
Covertly improving intelligibility with data-driven adaptations of speech timing
par: Tuttösí, Paige, et autres
Publié: (2026)
par: Tuttösí, Paige, et autres
Publié: (2026)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
par: Peng, Puyuan, et autres
Publié: (2025)
par: Peng, Puyuan, et autres
Publié: (2025)
BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition
par: Tuttösí, Paige, et autres
Publié: (2025)
par: Tuttösí, Paige, et autres
Publié: (2025)
SponTTS: modeling and transferring spontaneous style for TTS
par: Li, Hanzhao, et autres
Publié: (2023)
par: Li, Hanzhao, et autres
Publié: (2023)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
par: Nguyen, Tan Dat, et autres
Publié: (2025)
par: Nguyen, Tan Dat, et autres
Publié: (2025)
E1 TTS: Simple and Fast Non-Autoregressive TTS
par: Liu, Zhijun, et autres
Publié: (2024)
par: Liu, Zhijun, et autres
Publié: (2024)
Application of ASV for Voice Identification after VC and Duration Predictor Improvement in TTS Models
par: Nikolayevich, Borodin Kirill, et autres
Publié: (2024)
par: Nikolayevich, Borodin Kirill, et autres
Publié: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
par: Eskimez, Sefik Emre, et autres
Publié: (2024)
par: Eskimez, Sefik Emre, et autres
Publié: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
par: Qharabagh, Mahta Fetrat, et autres
Publié: (2024)
par: Qharabagh, Mahta Fetrat, et autres
Publié: (2024)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
par: Xue, Heyang, et autres
Publié: (2025)
par: Xue, Heyang, et autres
Publié: (2025)
Accent-VITS:accent transfer for end-to-end TTS
par: Ma, Linhan, et autres
Publié: (2023)
par: Ma, Linhan, et autres
Publié: (2023)
A Dataset for Automatic Assessment of TTS Quality in Spanish
par: Welford, Alejandro Sosa, et autres
Publié: (2025)
par: Welford, Alejandro Sosa, et autres
Publié: (2025)
Intelli-Z: Toward Intelligible Zero-Shot TTS
par: Jung, Sunghee, et autres
Publié: (2024)
par: Jung, Sunghee, et autres
Publié: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
par: Biadsy, Fadi, et autres
Publié: (2024)
par: Biadsy, Fadi, et autres
Publié: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
par: Liu, Huadai, et autres
Publié: (2023)
par: Liu, Huadai, et autres
Publié: (2023)
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
par: Zeldes, Ella, et autres
Publié: (2024)
par: Zeldes, Ella, et autres
Publié: (2024)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
par: He, Xinlu, et autres
Publié: (2025)
par: He, Xinlu, et autres
Publié: (2025)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
par: Li, Haoxun, et autres
Publié: (2025)
par: Li, Haoxun, et autres
Publié: (2025)
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
par: Cho, Chanhee, et autres
Publié: (2026)
par: Cho, Chanhee, et autres
Publié: (2026)
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
par: Hsu, Po-chun, et autres
Publié: (2023)
par: Hsu, Po-chun, et autres
Publié: (2023)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
par: Jeon, Yejin, et autres
Publié: (2024)
par: Jeon, Yejin, et autres
Publié: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
par: Ren, Yong, et autres
Publié: (2026)
par: Ren, Yong, et autres
Publié: (2026)
Bridging the gap between training and inference in LM-based TTS models
par: Zhang, Ruonan, et autres
Publié: (2025)
par: Zhang, Ruonan, et autres
Publié: (2025)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
par: Chary, Podakanti Satyajith
Publié: (2024)
par: Chary, Podakanti Satyajith
Publié: (2024)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
par: Zhong, Yi, et autres
Publié: (2023)
par: Zhong, Yi, et autres
Publié: (2023)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
par: Jiang, Ziyue, et autres
Publié: (2023)
par: Jiang, Ziyue, et autres
Publié: (2023)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
par: Guo, Hao-Han, et autres
Publié: (2025)
par: Guo, Hao-Han, et autres
Publié: (2025)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
par: Anastassiou, Philip, et autres
Publié: (2024)
par: Anastassiou, Philip, et autres
Publié: (2024)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
par: Wang, Jianzong, et autres
Publié: (2023)
par: Wang, Jianzong, et autres
Publié: (2023)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
par: Lu, Ye-Xin, et autres
Publié: (2025)
par: Lu, Ye-Xin, et autres
Publié: (2025)
Phone Duration Modeling for Speaker Age Estimation in Children
par: Shivakumar, Prashanth Gurunath, et autres
Publié: (2021)
par: Shivakumar, Prashanth Gurunath, et autres
Publié: (2021)
EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
par: Manku, Ruskin Raj, et autres
Publié: (2025)
par: Manku, Ruskin Raj, et autres
Publié: (2025)
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
par: Zhang, Jiawei, et autres
Publié: (2024)
par: Zhang, Jiawei, et autres
Publié: (2024)
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
par: Borodin, Kirill, et autres
Publié: (2026)
par: Borodin, Kirill, et autres
Publié: (2026)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
par: Ueda, Lucas H., et autres
Publié: (2024)
par: Ueda, Lucas H., et autres
Publié: (2024)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
par: Gudmalwar, Ashishkumar, et autres
Publié: (2024)
par: Gudmalwar, Ashishkumar, et autres
Publié: (2024)
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
par: Xie, Kun, et autres
Publié: (2025)
par: Xie, Kun, et autres
Publié: (2025)
Documents similaires
-
Mmm whatcha say? Uncovering distal and proximal context effects in first and second-language word perception using psychophysical reverse correlation
par: Tuttösí, Paige, et autres
Publié: (2024) -
I Know You're Listening: Adaptive Voice for HRI
par: Tuttösí, Paige
Publié: (2025) -
Quantification of Tenseness in English and Japanese Tense-Lax Vowels: A Lagrangian Model with Indicator θ1 and Force of Tenseness Ftense(t)
par: Ishizaki, Tatsuya
Publié: (2025) -
Covertly improving intelligibility with data-driven adaptations of speech timing
par: Tuttösí, Paige, et autres
Publié: (2026) -
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
par: Peng, Puyuan, et autres
Publié: (2025)