Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Tianle, Sun, Chengzhe, Rose, Phil, Jacobs, Cassandra L., Lyu, Siwei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Forensic deepfake audio detection using segmental speech features
por: Yang, Tianle, et al.
Publicado: (2025)
por: Yang, Tianle, et al.
Publicado: (2025)
Acoustic and perceptual differences between standard and accented speech and their voice clones
por: Yang, Tianle, et al.
Publicado: (2026)
por: Yang, Tianle, et al.
Publicado: (2026)
MOSS-TTS Technical Report
por: Gong, Yitian, et al.
Publicado: (2026)
por: Gong, Yitian, et al.
Publicado: (2026)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
por: Bataev, Vladimir, et al.
Publicado: (2025)
por: Bataev, Vladimir, et al.
Publicado: (2025)
A2TTS: TTS for Low Resource Indian Languages
por: Bhadoriya, Ayush Singh, et al.
Publicado: (2025)
por: Bhadoriya, Ayush Singh, et al.
Publicado: (2025)
Calliope: A TTS-based Narrated E-book Creator Ensuring Exact Synchronization, Privacy, and Layout Fidelity
por: Hammer, Hugo L., et al.
Publicado: (2026)
por: Hammer, Hugo L., et al.
Publicado: (2026)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
por: Lian, Jiachen, et al.
Publicado: (2022)
por: Lian, Jiachen, et al.
Publicado: (2022)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
por: Zhong, Yi, et al.
Publicado: (2023)
por: Zhong, Yi, et al.
Publicado: (2023)
Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
por: Shi, Yemin, et al.
Publicado: (2025)
por: Shi, Yemin, et al.
Publicado: (2025)
TTS-1 Technical Report
por: Atamanenko, Oleg, et al.
Publicado: (2025)
por: Atamanenko, Oleg, et al.
Publicado: (2025)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
por: Liu, Jiaxuan, et al.
Publicado: (2024)
por: Liu, Jiaxuan, et al.
Publicado: (2024)
RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning
por: Sigurgeirsson, Atli, et al.
Publicado: (2025)
por: Sigurgeirsson, Atli, et al.
Publicado: (2025)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
por: Shin, Seungyoun, et al.
Publicado: (2025)
por: Shin, Seungyoun, et al.
Publicado: (2025)
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
por: Chou, Cheng-Kang, et al.
Publicado: (2025)
por: Chou, Cheng-Kang, et al.
Publicado: (2025)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
por: Ma, Ziyang, et al.
Publicado: (2023)
por: Ma, Ziyang, et al.
Publicado: (2023)
Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation
por: He, Jiaxu, et al.
Publicado: (2026)
por: He, Jiaxu, et al.
Publicado: (2026)
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID
por: Handoyo, Ahmad Alfani, et al.
Publicado: (2024)
por: Handoyo, Ahmad Alfani, et al.
Publicado: (2024)
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
por: Zhao, Junchuan, et al.
Publicado: (2025)
por: Zhao, Junchuan, et al.
Publicado: (2025)
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
por: Xu, Ke, et al.
Publicado: (2026)
por: Xu, Ke, et al.
Publicado: (2026)
MunTTS: A Text-to-Speech System for Mundari
por: Gumma, Varun, et al.
Publicado: (2024)
por: Gumma, Varun, et al.
Publicado: (2024)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
ALICE: A Multifaceted Evaluation Framework of Large Audio-Language Models' In-Context Learning Ability
por: Piao, Yen-Ting, et al.
Publicado: (2026)
por: Piao, Yen-Ting, et al.
Publicado: (2026)
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
por: Deng, Yayue, et al.
Publicado: (2025)
por: Deng, Yayue, et al.
Publicado: (2025)
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
por: Zhou, Siyi, et al.
Publicado: (2025)
por: Zhou, Siyi, et al.
Publicado: (2025)
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
por: Song, Xingchen, et al.
Publicado: (2024)
por: Song, Xingchen, et al.
Publicado: (2024)
Robust TTS Training via Self-Purifying Flow Matching for the WildSpoof 2026 TTS Track
por: Yi, June Young, et al.
Publicado: (2025)
por: Yi, June Young, et al.
Publicado: (2025)
Assessing the Impact of Anisotropy in Neural Representations of Speech: A Case Study on Keyword Spotting
por: Wisniewski, Guillaume, et al.
Publicado: (2025)
por: Wisniewski, Guillaume, et al.
Publicado: (2025)
The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail
por: Menta, Venkata Pushpak Teja
Publicado: (2026)
por: Menta, Venkata Pushpak Teja
Publicado: (2026)
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
por: Wang, Rui, et al.
Publicado: (2025)
por: Wang, Rui, et al.
Publicado: (2025)
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
por: Liu, Yutong, et al.
Publicado: (2025)
por: Liu, Yutong, et al.
Publicado: (2025)
Assessing Latency in ASR Systems: A Methodological Perspective for Real-Time Use
por: Arriaga, Carlos, et al.
Publicado: (2024)
por: Arriaga, Carlos, et al.
Publicado: (2024)
Qwen3-TTS Technical Report
por: Hu, Hangrui, et al.
Publicado: (2026)
por: Hu, Hangrui, et al.
Publicado: (2026)
RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer
por: Matiyali, Neeraj, et al.
Publicado: (2025)
por: Matiyali, Neeraj, et al.
Publicado: (2025)
IndexTTS 2.5 Technical Report
por: Li, Yunpei, et al.
Publicado: (2026)
por: Li, Yunpei, et al.
Publicado: (2026)
Efficient Training for Cross-lingual Speech Language Models
por: Zhou, Yan, et al.
Publicado: (2026)
por: Zhou, Yan, et al.
Publicado: (2026)
PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech
por: Rahman, Hanif
Publicado: (2026)
por: Rahman, Hanif
Publicado: (2026)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
por: Roth, Amit, et al.
Publicado: (2024)
por: Roth, Amit, et al.
Publicado: (2024)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
por: Lee, Keon, et al.
Publicado: (2024)
por: Lee, Keon, et al.
Publicado: (2024)
Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum
por: Al-Radhi, Mohammed Salah, et al.
Publicado: (2026)
por: Al-Radhi, Mohammed Salah, et al.
Publicado: (2026)
Ejemplares similares
-
Forensic deepfake audio detection using segmental speech features
por: Yang, Tianle, et al.
Publicado: (2025) -
Acoustic and perceptual differences between standard and accented speech and their voice clones
por: Yang, Tianle, et al.
Publicado: (2026) -
MOSS-TTS Technical Report
por: Gong, Yitian, et al.
Publicado: (2026) -
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
por: Bataev, Vladimir, et al.
Publicado: (2025) -
A2TTS: TTS for Low Resource Indian Languages
por: Bhadoriya, Ayush Singh, et al.
Publicado: (2025)