GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zehua Kcriss, Chen, Meiying Melissa, Zhong, Yi, Liu, Pinxin, Duan, Zhiyao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
Generative Expressive Conversational Speech Synthesis
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
Deep Speech Synthesis from Multimodal Articulatory Representations
von: Wu, Peter, et al.
Veröffentlicht: (2024)
von: Wu, Peter, et al.
Veröffentlicht: (2024)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
von: Liu, Sen, et al.
Veröffentlicht: (2024)
von: Liu, Sen, et al.
Veröffentlicht: (2024)
Tracking Articulatory Dynamics in Speech with a Fixed-Weight BiLSTM-CNN Architecture
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
Acoustic to Articulatory Inversion of Speech; Data Driven Approaches, Challenges, Applications, and Future Scope
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2026)
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2026)
Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
von: Zhang, You, et al.
Veröffentlicht: (2021)
von: Zhang, You, et al.
Veröffentlicht: (2021)
Computational Narrative Understanding for Expressive Text-to-Speech
von: Michel, Gaspard, et al.
Veröffentlicht: (2025)
von: Michel, Gaspard, et al.
Veröffentlicht: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
(SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test
von: Bleeck, Stefan
Veröffentlicht: (2025)
von: Bleeck, Stefan
Veröffentlicht: (2025)
UR Channel-Robust Synthetic Speech Detection System for ASVspoof 2021
von: Chen, Xinhui, et al.
Veröffentlicht: (2021)
von: Chen, Xinhui, et al.
Veröffentlicht: (2021)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
von: Zhang, You, et al.
Veröffentlicht: (2025)
von: Zhang, You, et al.
Veröffentlicht: (2025)
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
von: Kang, Wonjune, et al.
Veröffentlicht: (2025)
von: Kang, Wonjune, et al.
Veröffentlicht: (2025)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
PhiNet: Speaker Verification with Phonetic Interpretability
von: Ma, Yi, et al.
Veröffentlicht: (2026)
von: Ma, Yi, et al.
Veröffentlicht: (2026)
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
von: Song, Yaodong, et al.
Veröffentlicht: (2025)
von: Song, Yaodong, et al.
Veröffentlicht: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
von: Han, Seungu, et al.
Veröffentlicht: (2026)
von: Han, Seungu, et al.
Veröffentlicht: (2026)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
On the Impact of Voice Anonymization on Speech Diagnostic Applications: a Case Study on COVID-19 Detection
von: Zhu, Yi, et al.
Veröffentlicht: (2023)
von: Zhu, Yi, et al.
Veröffentlicht: (2023)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
Multimodal Input Aids a Bayesian Model of Phonetic Learning
von: Zhi, Sophia, et al.
Veröffentlicht: (2024)
von: Zhi, Sophia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
von: Chen, Meiying, et al.
Veröffentlicht: (2022) -
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025) -
Generative Expressive Conversational Speech Synthesis
von: Liu, Rui, et al.
Veröffentlicht: (2024) -
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024) -
Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
von: Zhang, Yu, et al.
Veröffentlicht: (2025)