FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912354232434688 |
|---|---|
| author | Li, Hanzhao Li, Yuke Wang, Xinsheng Hu, Jingbin Xie, Qicong Yang, Shan Xie, Lei |
| author_facet | Li, Hanzhao Li, Yuke Wang, Xinsheng Hu, Jingbin Xie, Qicong Yang, Shan Xie, Lei |
| contents | Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjusting the style while preserving a selected speaker's timbre, or choosing a style and generating a voice that matches a character's visual appearance. To overcome these challenges, we propose \textit{FleSpeech}, a novel multi-stage speech generation framework that allows for more flexible manipulation of speech attributes by integrating various forms of control. FleSpeech employs a multimodal prompt encoder that processes and unifies different text, audio, and visual prompts into a cohesive representation. This approach enhances the adaptability of speech synthesis and supports creative and precise control over the generated speech. Additionally, we develop a data collection pipeline for multimodal datasets to facilitate further research and applications in this field. Comprehensive subjective and objective experiments demonstrate the effectiveness of FleSpeech. Audio samples are available at https://kkksuper.github.io/FleSpeech/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_04644 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Li, Hanzhao Li, Yuke Wang, Xinsheng Hu, Jingbin Xie, Qicong Yang, Shan Xie, Lei Audio and Speech Processing Sound Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjusting the style while preserving a selected speaker's timbre, or choosing a style and generating a voice that matches a character's visual appearance. To overcome these challenges, we propose \textit{FleSpeech}, a novel multi-stage speech generation framework that allows for more flexible manipulation of speech attributes by integrating various forms of control. FleSpeech employs a multimodal prompt encoder that processes and unifies different text, audio, and visual prompts into a cohesive representation. This approach enhances the adaptability of speech synthesis and supports creative and precise control over the generated speech. Additionally, we develop a data collection pipeline for multimodal datasets to facilitate further research and applications in this field. Comprehensive subjective and objective experiments demonstrate the effectiveness of FleSpeech. Audio samples are available at https://kkksuper.github.io/FleSpeech/ |
| title | FleSpeech: Flexibly Controllable Speech Generation with Various Prompts |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2501.04644 |