Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Tianle, Sun, Chengzhe, Rose, Phil, Jacobs, Cassandra L., Lyu, Siwei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Forensic deepfake audio detection using segmental speech features
di: Yang, Tianle, et al.
Pubblicazione: (2025)
di: Yang, Tianle, et al.
Pubblicazione: (2025)
Acoustic and perceptual differences between standard and accented speech and their voice clones
di: Yang, Tianle, et al.
Pubblicazione: (2026)
di: Yang, Tianle, et al.
Pubblicazione: (2026)
MOSS-TTS Technical Report
di: Gong, Yitian, et al.
Pubblicazione: (2026)
di: Gong, Yitian, et al.
Pubblicazione: (2026)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
A2TTS: TTS for Low Resource Indian Languages
di: Bhadoriya, Ayush Singh, et al.
Pubblicazione: (2025)
di: Bhadoriya, Ayush Singh, et al.
Pubblicazione: (2025)
Calliope: A TTS-based Narrated E-book Creator Ensuring Exact Synchronization, Privacy, and Layout Fidelity
di: Hammer, Hugo L., et al.
Pubblicazione: (2026)
di: Hammer, Hugo L., et al.
Pubblicazione: (2026)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
di: Lian, Jiachen, et al.
Pubblicazione: (2022)
di: Lian, Jiachen, et al.
Pubblicazione: (2022)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
di: Zhong, Yi, et al.
Pubblicazione: (2023)
di: Zhong, Yi, et al.
Pubblicazione: (2023)
Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
di: Shi, Yemin, et al.
Pubblicazione: (2025)
di: Shi, Yemin, et al.
Pubblicazione: (2025)
TTS-1 Technical Report
di: Atamanenko, Oleg, et al.
Pubblicazione: (2025)
di: Atamanenko, Oleg, et al.
Pubblicazione: (2025)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning
di: Sigurgeirsson, Atli, et al.
Pubblicazione: (2025)
di: Sigurgeirsson, Atli, et al.
Pubblicazione: (2025)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
di: Shin, Seungyoun, et al.
Pubblicazione: (2025)
di: Shin, Seungyoun, et al.
Pubblicazione: (2025)
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
di: Chou, Cheng-Kang, et al.
Pubblicazione: (2025)
di: Chou, Cheng-Kang, et al.
Pubblicazione: (2025)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation
di: He, Jiaxu, et al.
Pubblicazione: (2026)
di: He, Jiaxu, et al.
Pubblicazione: (2026)
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID
di: Handoyo, Ahmad Alfani, et al.
Pubblicazione: (2024)
di: Handoyo, Ahmad Alfani, et al.
Pubblicazione: (2024)
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
di: Zhao, Junchuan, et al.
Pubblicazione: (2025)
di: Zhao, Junchuan, et al.
Pubblicazione: (2025)
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
di: Xu, Ke, et al.
Pubblicazione: (2026)
di: Xu, Ke, et al.
Pubblicazione: (2026)
MunTTS: A Text-to-Speech System for Mundari
di: Gumma, Varun, et al.
Pubblicazione: (2024)
di: Gumma, Varun, et al.
Pubblicazione: (2024)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
di: Ghosh, Sreyan, et al.
Pubblicazione: (2024)
ALICE: A Multifaceted Evaluation Framework of Large Audio-Language Models' In-Context Learning Ability
di: Piao, Yen-Ting, et al.
Pubblicazione: (2026)
di: Piao, Yen-Ting, et al.
Pubblicazione: (2026)
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
di: Deng, Yayue, et al.
Pubblicazione: (2025)
di: Deng, Yayue, et al.
Pubblicazione: (2025)
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
di: Zhou, Siyi, et al.
Pubblicazione: (2025)
di: Zhou, Siyi, et al.
Pubblicazione: (2025)
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
di: Song, Xingchen, et al.
Pubblicazione: (2024)
di: Song, Xingchen, et al.
Pubblicazione: (2024)
Robust TTS Training via Self-Purifying Flow Matching for the WildSpoof 2026 TTS Track
di: Yi, June Young, et al.
Pubblicazione: (2025)
di: Yi, June Young, et al.
Pubblicazione: (2025)
Assessing the Impact of Anisotropy in Neural Representations of Speech: A Case Study on Keyword Spotting
di: Wisniewski, Guillaume, et al.
Pubblicazione: (2025)
di: Wisniewski, Guillaume, et al.
Pubblicazione: (2025)
The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail
di: Menta, Venkata Pushpak Teja
Pubblicazione: (2026)
di: Menta, Venkata Pushpak Teja
Pubblicazione: (2026)
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
di: Wang, Rui, et al.
Pubblicazione: (2025)
di: Wang, Rui, et al.
Pubblicazione: (2025)
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
di: Liu, Yutong, et al.
Pubblicazione: (2025)
di: Liu, Yutong, et al.
Pubblicazione: (2025)
Assessing Latency in ASR Systems: A Methodological Perspective for Real-Time Use
di: Arriaga, Carlos, et al.
Pubblicazione: (2024)
di: Arriaga, Carlos, et al.
Pubblicazione: (2024)
Qwen3-TTS Technical Report
di: Hu, Hangrui, et al.
Pubblicazione: (2026)
di: Hu, Hangrui, et al.
Pubblicazione: (2026)
RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer
di: Matiyali, Neeraj, et al.
Pubblicazione: (2025)
di: Matiyali, Neeraj, et al.
Pubblicazione: (2025)
IndexTTS 2.5 Technical Report
di: Li, Yunpei, et al.
Pubblicazione: (2026)
di: Li, Yunpei, et al.
Pubblicazione: (2026)
Efficient Training for Cross-lingual Speech Language Models
di: Zhou, Yan, et al.
Pubblicazione: (2026)
di: Zhou, Yan, et al.
Pubblicazione: (2026)
PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech
di: Rahman, Hanif
Pubblicazione: (2026)
di: Rahman, Hanif
Pubblicazione: (2026)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
di: Roth, Amit, et al.
Pubblicazione: (2024)
di: Roth, Amit, et al.
Pubblicazione: (2024)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
di: Lee, Keon, et al.
Pubblicazione: (2024)
di: Lee, Keon, et al.
Pubblicazione: (2024)
Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum
di: Al-Radhi, Mohammed Salah, et al.
Pubblicazione: (2026)
di: Al-Radhi, Mohammed Salah, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Forensic deepfake audio detection using segmental speech features
di: Yang, Tianle, et al.
Pubblicazione: (2025) -
Acoustic and perceptual differences between standard and accented speech and their voice clones
di: Yang, Tianle, et al.
Pubblicazione: (2026) -
MOSS-TTS Technical Report
di: Gong, Yitian, et al.
Pubblicazione: (2026) -
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
di: Bataev, Vladimir, et al.
Pubblicazione: (2025) -
A2TTS: TTS for Low Resource Indian Languages
di: Bhadoriya, Ayush Singh, et al.
Pubblicazione: (2025)