GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Yingying, Zhang, Shilei, Deng, Chao, Feng, Junlan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
di: Yang, Runyan, et al.
Pubblicazione: (2025)
di: Yang, Runyan, et al.
Pubblicazione: (2025)
HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
di: Si, Yuke, et al.
Pubblicazione: (2025)
di: Si, Yuke, et al.
Pubblicazione: (2025)
MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
di: Sun, Haiyang, et al.
Pubblicazione: (2023)
di: Sun, Haiyang, et al.
Pubblicazione: (2023)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
di: Yang, Runyan, et al.
Pubblicazione: (2024)
di: Yang, Runyan, et al.
Pubblicazione: (2024)
On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations
di: Hao, Yaqian, et al.
Pubblicazione: (2024)
di: Hao, Yaqian, et al.
Pubblicazione: (2024)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
di: Wang, Chen, et al.
Pubblicazione: (2024)
di: Wang, Chen, et al.
Pubblicazione: (2024)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
di: Chen, Yanan, et al.
Pubblicazione: (2024)
di: Chen, Yanan, et al.
Pubblicazione: (2024)
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
di: Zhu, Yongxin, et al.
Pubblicazione: (2024)
di: Zhu, Yongxin, et al.
Pubblicazione: (2024)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
di: Ferraz, Thomas Palmeira, et al.
Pubblicazione: (2023)
di: Ferraz, Thomas Palmeira, et al.
Pubblicazione: (2023)
To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
Exploring Energy-Based Models for Out-of-Distribution Detection in Dialect Identification
di: Hao, Yaqian, et al.
Pubblicazione: (2024)
di: Hao, Yaqian, et al.
Pubblicazione: (2024)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
CEC: A Noisy Label Detection Method for Speaker Recognition
di: Shen, Yao, et al.
Pubblicazione: (2024)
di: Shen, Yao, et al.
Pubblicazione: (2024)
Efficient Interleaved Speech Modeling through Knowledge Distillation
di: Nouriborji, Mohammadmahdi, et al.
Pubblicazione: (2025)
di: Nouriborji, Mohammadmahdi, et al.
Pubblicazione: (2025)
Efficient Speech Translation through Model Compression and Knowledge Distillation
di: Moslem, Yasmin
Pubblicazione: (2025)
di: Moslem, Yasmin
Pubblicazione: (2025)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
di: Wang, Yujin, et al.
Pubblicazione: (2022)
di: Wang, Yujin, et al.
Pubblicazione: (2022)
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers
di: Hentschel, Michael, et al.
Pubblicazione: (2024)
di: Hentschel, Michael, et al.
Pubblicazione: (2024)
SONAR: Self-Distilled Continual Pre-training for Domain Adaptive Audio Representation
di: Zhang, Yizhou, et al.
Pubblicazione: (2025)
di: Zhang, Yizhou, et al.
Pubblicazione: (2025)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
di: Liu, Alexander H., et al.
Pubblicazione: (2025)
di: Liu, Alexander H., et al.
Pubblicazione: (2025)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2026)
di: Wang, Zhichao, et al.
Pubblicazione: (2026)
Guided by the Plan: Enhancing Faithful Autoregressive Text-to-Audio Generation with Guided Decoding
di: Wang, Juncheng, et al.
Pubblicazione: (2026)
di: Wang, Juncheng, et al.
Pubblicazione: (2026)
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
di: Bijoy, Mehedi Hasan, et al.
Pubblicazione: (2025)
di: Bijoy, Mehedi Hasan, et al.
Pubblicazione: (2025)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
di: Hwang, Min-Jae, et al.
Pubblicazione: (2024)
di: Hwang, Min-Jae, et al.
Pubblicazione: (2024)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
di: Jang, Kangwook, et al.
Pubblicazione: (2023)
di: Jang, Kangwook, et al.
Pubblicazione: (2023)
BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
di: Wang, Chen, et al.
Pubblicazione: (2023)
di: Wang, Chen, et al.
Pubblicazione: (2023)
Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study
di: Huang, W. Ronny, et al.
Pubblicazione: (2024)
di: Huang, W. Ronny, et al.
Pubblicazione: (2024)
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
di: Benita, Roi, et al.
Pubblicazione: (2023)
di: Benita, Roi, et al.
Pubblicazione: (2023)
tinyCLAP: Distilling Constrastive Language-Audio Pretrained Models
di: Paissan, Francesco, et al.
Pubblicazione: (2023)
di: Paissan, Francesco, et al.
Pubblicazione: (2023)
Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
di: Dong, Lukuan, et al.
Pubblicazione: (2024)
di: Dong, Lukuan, et al.
Pubblicazione: (2024)
Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
di: Pan, Zihan, et al.
Pubblicazione: (2024)
di: Pan, Zihan, et al.
Pubblicazione: (2024)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
di: Xue, Jinlong, et al.
Pubblicazione: (2024)
di: Xue, Jinlong, et al.
Pubblicazione: (2024)
Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2024)
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2024)
USAD: Universal Speech and Audio Representation via Distillation
di: Chang, Heng-Jui, et al.
Pubblicazione: (2025)
di: Chang, Heng-Jui, et al.
Pubblicazione: (2025)
SALMONN: Towards Generic Hearing Abilities for Large Language Models
di: Tang, Changli, et al.
Pubblicazione: (2023)
di: Tang, Changli, et al.
Pubblicazione: (2023)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
di: Zeng, Aohan, et al.
Pubblicazione: (2024)
di: Zeng, Aohan, et al.
Pubblicazione: (2024)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
Fine-grained Speech Sentiment Analysis in Chinese Psychological Support Hotlines Based on Large-scale Pre-trained Model
di: Chen, Zhonglong, et al.
Pubblicazione: (2024)
di: Chen, Zhonglong, et al.
Pubblicazione: (2024)
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
di: Lehečka, Jan, et al.
Pubblicazione: (2024)
di: Lehečka, Jan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
di: Yang, Runyan, et al.
Pubblicazione: (2025) -
HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
di: Si, Yuke, et al.
Pubblicazione: (2025) -
MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
di: Sun, Haiyang, et al.
Pubblicazione: (2023) -
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
di: Yang, Runyan, et al.
Pubblicazione: (2024) -
On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations
di: Hao, Yaqian, et al.
Pubblicazione: (2024)