SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen, Tan Dat, Kim, Jaehun, Kim, Ji-Hoon, Choi, Shukjae, Lim, Youshin, Chung, Joon Son |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaptVC: High Quality Voice Conversion with Adaptive Learning
von: Kim, Jaehun, et al.
Veröffentlicht: (2025)
von: Kim, Jaehun, et al.
Veröffentlicht: (2025)
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
Boosting Unknown-number Speaker Separation with Transformer Decoder-based Attractor
von: Lee, Younglo, et al.
Veröffentlicht: (2024)
von: Lee, Younglo, et al.
Veröffentlicht: (2024)
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2026)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2026)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
Lightweight Audio Segmentation for Long-form Speech Translation
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
von: Cho, Chanhee, et al.
Veröffentlicht: (2026)
von: Cho, Chanhee, et al.
Veröffentlicht: (2026)
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
von: Choi, Youngwon, et al.
Veröffentlicht: (2026)
von: Choi, Youngwon, et al.
Veröffentlicht: (2026)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
DISPATCH: Distilling Selective Patches for Speech Enhancement
von: Kim, Dohwan, et al.
Veröffentlicht: (2025)
von: Kim, Dohwan, et al.
Veröffentlicht: (2025)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
Adversarial Training of Denoising Diffusion Model Using Dual Discriminators for High-Fidelity Multi-Speaker TTS
von: Ko, Myeongjin, et al.
Veröffentlicht: (2023)
von: Ko, Myeongjin, et al.
Veröffentlicht: (2023)
Intelli-Z: Toward Intelligible Zero-Shot TTS
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning
von: Han, Bing, et al.
Veröffentlicht: (2024)
von: Han, Bing, et al.
Veröffentlicht: (2024)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
From Coarse to Fine: Efficient Training for Audio Spectrogram Transformers
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
E1 TTS: Simple and Fast Non-Autoregressive TTS
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
Mamba2 Meets Silence: Robust Vocal Source Separation for Sparse Regions
von: Kim, Euiyeon, et al.
Veröffentlicht: (2025)
von: Kim, Euiyeon, et al.
Veröffentlicht: (2025)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AdaptVC: High Quality Voice Conversion with Adaptive Learning
von: Kim, Jaehun, et al.
Veröffentlicht: (2025) -
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024) -
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023) -
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025) -
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)