DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huo, Yanru, Jiang, Ziyue, Tang, Zuoli, Hong, Qingyang, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
von: Wang, Jinting, et al.
Veröffentlicht: (2025)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
Score-Based Training for Energy-Based TTS Models
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
E1 TTS: Simple and Fast Non-Autoregressive TTS
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
von: Xin, Detai, et al.
Veröffentlicht: (2026)
von: Xin, Detai, et al.
Veröffentlicht: (2026)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
Qwen3-TTS Technical Report
von: Hu, Hangrui, et al.
Veröffentlicht: (2026)
von: Hu, Hangrui, et al.
Veröffentlicht: (2026)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
von: Xie, Tianxin, et al.
Veröffentlicht: (2025)
von: Xie, Tianxin, et al.
Veröffentlicht: (2025)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
von: Jia, Yuhang, et al.
Veröffentlicht: (2024)
von: Jia, Yuhang, et al.
Veröffentlicht: (2024)
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
von: Chen, Peijie, et al.
Veröffentlicht: (2025)
von: Chen, Peijie, et al.
Veröffentlicht: (2025)
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
von: Xie, Kun, et al.
Veröffentlicht: (2025)
von: Xie, Kun, et al.
Veröffentlicht: (2025)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2025)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2025)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
A Dataset for Automatic Assessment of TTS Quality in Spanish
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
Intelli-Z: Toward Intelligible Zero-Shot TTS
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS
von: Seo, Deokjin, et al.
Veröffentlicht: (2026)
von: Seo, Deokjin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
von: Xue, Heyang, et al.
Veröffentlicht: (2025) -
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
von: Guan, Wenhao, et al.
Veröffentlicht: (2025) -
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
von: Guan, Wenhao, et al.
Veröffentlicht: (2023) -
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
von: Cao, Tianyu, et al.
Veröffentlicht: (2026) -
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
von: Wang, Jinting, et al.
Veröffentlicht: (2025)