Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
Fuente:
arXiv
Guardado en:
| Autores principales: | C, Shiva Kumar, Dhiman, Jitendra Kumar, Adiga, Nagaraj, Singh, Shatrughan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
por: Stahl, Benjamin, et al.
Publicado: (2025)
por: Stahl, Benjamin, et al.
Publicado: (2025)
Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing
por: Peng, Junyi, et al.
Publicado: (2025)
por: Peng, Junyi, et al.
Publicado: (2025)
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
por: Han, Jiangyu, et al.
Publicado: (2025)
por: Han, Jiangyu, et al.
Publicado: (2025)
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
por: Han, Jiangyu, et al.
Publicado: (2025)
por: Han, Jiangyu, et al.
Publicado: (2025)
Rethinking Mamba in Speech Processing by Self-Supervised Models
por: Zhang, Xiangyu, et al.
Publicado: (2024)
por: Zhang, Xiangyu, et al.
Publicado: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
por: Wang, Yujin, et al.
Publicado: (2022)
por: Wang, Yujin, et al.
Publicado: (2022)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
por: Cohen, Eyal, et al.
Publicado: (2025)
por: Cohen, Eyal, et al.
Publicado: (2025)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
por: Li, Jialu, et al.
Publicado: (2024)
por: Li, Jialu, et al.
Publicado: (2024)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
por: Jang, Kangwook, et al.
Publicado: (2023)
por: Jang, Kangwook, et al.
Publicado: (2023)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
por: Hwang, Min-Jae, et al.
Publicado: (2024)
por: Hwang, Min-Jae, et al.
Publicado: (2024)
On the use of Performer and Agent Attention for Spoken Language Identification
por: dhiman, Jitendra Kumar, et al.
Publicado: (2025)
por: dhiman, Jitendra Kumar, et al.
Publicado: (2025)
Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
por: Li, Ze, et al.
Publicado: (2025)
por: Li, Ze, et al.
Publicado: (2025)
Context-Driven Dynamic Pruning for Large Speech Foundation Models
por: Someki, Masao, et al.
Publicado: (2025)
por: Someki, Masao, et al.
Publicado: (2025)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
por: Kumar, Shashi, et al.
Publicado: (2024)
por: Kumar, Shashi, et al.
Publicado: (2024)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
por: Nguyen, Tan Dat, et al.
Publicado: (2025)
por: Nguyen, Tan Dat, et al.
Publicado: (2025)
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
por: Guo, Xin, et al.
Publicado: (2026)
por: Guo, Xin, et al.
Publicado: (2026)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
por: Farhadipour, Aref, et al.
Publicado: (2024)
por: Farhadipour, Aref, et al.
Publicado: (2024)
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
por: Yuan, Xihao, et al.
Publicado: (2025)
por: Yuan, Xihao, et al.
Publicado: (2025)
Frequency-mix Knowledge Distillation for Fake Speech Detection
por: Fan, Cunhang, et al.
Publicado: (2024)
por: Fan, Cunhang, et al.
Publicado: (2024)
A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation
por: Zhang, En-Wei, et al.
Publicado: (2025)
por: Zhang, En-Wei, et al.
Publicado: (2025)
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
por: Girish, et al.
Publicado: (2026)
por: Girish, et al.
Publicado: (2026)
Multi-Distillation from Speech and Music Representation Models
por: Wei, Jui-Chiang, et al.
Publicado: (2025)
por: Wei, Jui-Chiang, et al.
Publicado: (2025)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
por: Ashihara, Takanori, et al.
Publicado: (2023)
por: Ashihara, Takanori, et al.
Publicado: (2023)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
por: Maekaku, Takashi, et al.
Publicado: (2025)
por: Maekaku, Takashi, et al.
Publicado: (2025)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
Cyclostationarity Analysis as a Complement to Self-Supervised Representations for Speech Deepfake Detection
por: Hanilçi, Cemal, et al.
Publicado: (2026)
por: Hanilçi, Cemal, et al.
Publicado: (2026)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
Leveraging Self-Supervised Audio-Visual Pretrained Models to Improve Vocoded Speech Intelligibility in Cochlear Implant Simulation
por: Lai, Richard Lee, et al.
Publicado: (2023)
por: Lai, Richard Lee, et al.
Publicado: (2023)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
por: Ogg, Mattson, et al.
Publicado: (2025)
por: Ogg, Mattson, et al.
Publicado: (2025)
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
por: Terashima, Ryo, et al.
Publicado: (2025)
por: Terashima, Ryo, et al.
Publicado: (2025)
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
por: Lin, Jingru, et al.
Publicado: (2024)
por: Lin, Jingru, et al.
Publicado: (2024)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
por: Fan, Cunhang, et al.
Publicado: (2023)
por: Fan, Cunhang, et al.
Publicado: (2023)
Task-Agnostic Structured Pruning of Speech Representation Models
por: Wang, Haoyu, et al.
Publicado: (2023)
por: Wang, Haoyu, et al.
Publicado: (2023)
Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
por: Han, Runduo, et al.
Publicado: (2024)
por: Han, Runduo, et al.
Publicado: (2024)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
por: Cui, Yang, et al.
Publicado: (2025)
por: Cui, Yang, et al.
Publicado: (2025)
SepPrune: Structured Pruning for Efficient Deep Speech Separation
por: Li, Yuqi, et al.
Publicado: (2025)
por: Li, Yuqi, et al.
Publicado: (2025)
Comparison of Knowledge Distillation Methods for Low-complexity Multi-microphone Speech Enhancement using the FT-JNF Architecture
por: Metzger, Robert, et al.
Publicado: (2025)
por: Metzger, Robert, et al.
Publicado: (2025)
Enhancing Fully Formatted End-to-End Speech Recognition with Knowledge Distillation via Multi-Codebook Vector Quantization
por: You, Jian, et al.
Publicado: (2025)
por: You, Jian, et al.
Publicado: (2025)
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
por: Yang, Yifan, et al.
Publicado: (2024)
por: Yang, Yifan, et al.
Publicado: (2024)
Ejemplares similares
-
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
por: Stahl, Benjamin, et al.
Publicado: (2025) -
Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing
por: Peng, Junyi, et al.
Publicado: (2025) -
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
por: Han, Jiangyu, et al.
Publicado: (2025) -
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
por: Han, Jiangyu, et al.
Publicado: (2025) -
Rethinking Mamba in Speech Processing by Self-Supervised Models
por: Zhang, Xiangyu, et al.
Publicado: (2024)