DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qi, Xin, Fu, Ruibo, Wen, Zhengqi, Wang, Tao, Qiang, Chunyu, Tao, Jianhua, Li, Chenxing, Lu, Yi, Shi, Shuchen, Wang, Zhiyong, Wang, Xiaopeng, Xie, Yuankun, Liu, Yukun, Liu, Xuefei, Li, Guanjun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910610184208384
author Qi, Xin
Fu, Ruibo
Wen, Zhengqi
Wang, Tao
Qiang, Chunyu
Tao, Jianhua
Li, Chenxing
Lu, Yi
Shi, Shuchen
Wang, Zhiyong
Wang, Xiaopeng
Xie, Yuankun
Liu, Yukun
Liu, Xuefei
Li, Guanjun
author_facet Qi, Xin
Fu, Ruibo
Wen, Zhengqi
Wang, Tao
Qiang, Chunyu
Tao, Jianhua
Li, Chenxing
Lu, Yi
Shi, Shuchen
Wang, Zhiyong
Wang, Xiaopeng
Xie, Yuankun
Liu, Yukun
Liu, Xuefei
Li, Guanjun
contents In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models.
format Preprint
id arxiv_https___arxiv_org_abs_2409_11835
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
Qi, Xin
Fu, Ruibo
Wen, Zhengqi
Wang, Tao
Qiang, Chunyu
Tao, Jianhua
Li, Chenxing
Lu, Yi
Shi, Shuchen
Wang, Zhiyong
Wang, Xiaopeng
Xie, Yuankun
Liu, Yukun
Liu, Xuefei
Li, Guanjun
Sound
Artificial Intelligence
Audio and Speech Processing
In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models.
title DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2409.11835