DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910610184208384 |
|---|---|
| author | Qi, Xin Fu, Ruibo Wen, Zhengqi Wang, Tao Qiang, Chunyu Tao, Jianhua Li, Chenxing Lu, Yi Shi, Shuchen Wang, Zhiyong Wang, Xiaopeng Xie, Yuankun Liu, Yukun Liu, Xuefei Li, Guanjun |
| author_facet | Qi, Xin Fu, Ruibo Wen, Zhengqi Wang, Tao Qiang, Chunyu Tao, Jianhua Li, Chenxing Lu, Yi Shi, Shuchen Wang, Zhiyong Wang, Xiaopeng Xie, Yuankun Liu, Yukun Liu, Xuefei Li, Guanjun |
| contents | In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_11835 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech Qi, Xin Fu, Ruibo Wen, Zhengqi Wang, Tao Qiang, Chunyu Tao, Jianhua Li, Chenxing Lu, Yi Shi, Shuchen Wang, Zhiyong Wang, Xiaopeng Xie, Yuankun Liu, Yukun Liu, Xuefei Li, Guanjun Sound Artificial Intelligence Audio and Speech Processing In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models. |
| title | DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech |
| topic | Sound Artificial Intelligence Audio and Speech Processing |
| url | https://arxiv.org/abs/2409.11835 |