DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huo, Yanru, Jiang, Ziyue, Tang, Zuoli, Hong, Qingyang, Zhao, Zhou
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916947217612800
author Huo, Yanru
Jiang, Ziyue
Tang, Zuoli
Hong, Qingyang
Zhao, Zhou
author_facet Huo, Yanru
Jiang, Ziyue
Tang, Zuoli
Hong, Qingyang
Zhao, Zhou
contents While Diffusion Transformers (DiT) have advanced non-autoregressive (NAR) speech synthesis, their high computational demands remain an limitation. Existing DiT-based text-to-speech (TTS) model acceleration approaches mainly focus on reducing sampling steps through distillation techniques, yet they remain constrained by training costs. We introduce DiTReducio, a training-free acceleration framework that compresses computations in DiT-based TTS models via progressive calibration. We propose two compression methods, Temporal Skipping and Branch Skipping, to eliminate redundant computations during inference. Moreover, based on two characteristic attention patterns identified within DiT layers, we devise a pattern-guided strategy to selectively apply the compression methods. Our method allows flexible modulation between generation quality and computational efficiency through adjustable compression thresholds. Experimental evaluations conducted on F5-TTS and MegaTTS 3 demonstrate that DiTReducio achieves a 75.4% reduction in FLOPs and improves the Real-Time Factor (RTF) by 37.1%, while preserving generation quality.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09748
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
Huo, Yanru
Jiang, Ziyue
Tang, Zuoli
Hong, Qingyang
Zhao, Zhou
Sound
Audio and Speech Processing
While Diffusion Transformers (DiT) have advanced non-autoregressive (NAR) speech synthesis, their high computational demands remain an limitation. Existing DiT-based text-to-speech (TTS) model acceleration approaches mainly focus on reducing sampling steps through distillation techniques, yet they remain constrained by training costs. We introduce DiTReducio, a training-free acceleration framework that compresses computations in DiT-based TTS models via progressive calibration. We propose two compression methods, Temporal Skipping and Branch Skipping, to eliminate redundant computations during inference. Moreover, based on two characteristic attention patterns identified within DiT layers, we devise a pattern-guided strategy to selectively apply the compression methods. Our method allows flexible modulation between generation quality and computational efficiency through adjustable compression thresholds. Experimental evaluations conducted on F5-TTS and MegaTTS 3 demonstrate that DiTReducio achieves a 75.4% reduction in FLOPs and improves the Real-Time Factor (RTF) by 37.1%, while preserving generation quality.
title DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.09748