Adaptive Knowledge Distillation for Device-Directed Speech Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | Chi, Hyung Gun, Pesce, Florian, Chang, Wonil, Rudovic, Oggi, Argueta, Arturo, Braun, Stefan, Garg, Vineet, Abdelaziz, Ahmed Hussen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
por: Chi, Hyung Gun, et al.
Publicado: (2025)
por: Chi, Hyung Gun, et al.
Publicado: (2025)
Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
por: Ognjen, et al.
Publicado: (2024)
por: Ognjen, et al.
Publicado: (2024)
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
por: Palaskar, Shruti, et al.
Publicado: (2024)
por: Palaskar, Shruti, et al.
Publicado: (2024)
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
por: Yuan, Xihao, et al.
Publicado: (2025)
por: Yuan, Xihao, et al.
Publicado: (2025)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
por: Chen, Li-Wei, et al.
Publicado: (2024)
por: Chen, Li-Wei, et al.
Publicado: (2024)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study
por: Brueggeman, Avamarie, et al.
Publicado: (2023)
por: Brueggeman, Avamarie, et al.
Publicado: (2023)
Frequency-mix Knowledge Distillation for Fake Speech Detection
por: Fan, Cunhang, et al.
Publicado: (2024)
por: Fan, Cunhang, et al.
Publicado: (2024)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
por: Fan, Cunhang, et al.
Publicado: (2023)
por: Fan, Cunhang, et al.
Publicado: (2023)
Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
por: Han, Runduo, et al.
Publicado: (2024)
por: Han, Runduo, et al.
Publicado: (2024)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
por: Cui, Yang, et al.
Publicado: (2025)
por: Cui, Yang, et al.
Publicado: (2025)
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
Leveraging Local and Global Knowledge Integration with Time-Frequency Calibrated Distillation for Speech Enhancement
por: Cheng, Jiaming, et al.
Publicado: (2025)
por: Cheng, Jiaming, et al.
Publicado: (2025)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
Ultra-Low Latency Speech Enhancement - A Comprehensive Study
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
Asymmetric Encoder-Decoder Based on Time-Frequency Correlation for Speech Separation
por: Shin, Ui-Hyeop, et al.
Publicado: (2026)
por: Shin, Ui-Hyeop, et al.
Publicado: (2026)
Statistical Beamformer Exploiting Non-stationarity and Sparsity with Spatially Constrained ICA for Robust Speech Recognition
por: Shin, Ui-Hyeop, et al.
Publicado: (2023)
por: Shin, Ui-Hyeop, et al.
Publicado: (2023)
DISPATCH: Distilling Selective Patches for Speech Enhancement
por: Kim, Dohwan, et al.
Publicado: (2025)
por: Kim, Dohwan, et al.
Publicado: (2025)
Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
por: Yang, Qingran, et al.
Publicado: (2026)
por: Yang, Qingran, et al.
Publicado: (2026)
Dataset-Distillation Generative Model for Speech Emotion Recognition
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2024)
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2024)
Efficient Interleaved Speech Modeling through Knowledge Distillation
por: Nouriborji, Mohammadmahdi, et al.
Publicado: (2025)
por: Nouriborji, Mohammadmahdi, et al.
Publicado: (2025)
TF-CorrNet: Leveraging Spatial Correlation for Continuous Speech Separation
por: Shin, Ui-Hyeop, et al.
Publicado: (2025)
por: Shin, Ui-Hyeop, et al.
Publicado: (2025)
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
por: Kounadis-Bastian, Dionyssos, et al.
Publicado: (2024)
por: Kounadis-Bastian, Dionyssos, et al.
Publicado: (2024)
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
por: Shah, Neil, et al.
Publicado: (2024)
por: Shah, Neil, et al.
Publicado: (2024)
Reverse Attention for Lightweight Speech Enhancement on Edge Devices
por: Ojha, Shuubham, et al.
Publicado: (2025)
por: Ojha, Shuubham, et al.
Publicado: (2025)
Robust One-step Speech Enhancement via Consistency Distillation
por: Xu, Liang, et al.
Publicado: (2025)
por: Xu, Liang, et al.
Publicado: (2025)
NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping
por: Büthe, Jan, et al.
Publicado: (2023)
por: Büthe, Jan, et al.
Publicado: (2023)
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
por: Jung, Jee-weon, et al.
Publicado: (2024)
por: Jung, Jee-weon, et al.
Publicado: (2024)
Efficient Speech Translation through Model Compression and Knowledge Distillation
por: Moslem, Yasmin
Publicado: (2025)
por: Moslem, Yasmin
Publicado: (2025)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
por: Serre, Thomas, et al.
Publicado: (2026)
por: Serre, Thomas, et al.
Publicado: (2026)
Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification
por: Truong, Duc-Tuan, et al.
Publicado: (2023)
por: Truong, Duc-Tuan, et al.
Publicado: (2023)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
por: Xu, Xuenan, et al.
Publicado: (2024)
por: Xu, Xuenan, et al.
Publicado: (2024)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
por: Wang, Anna, et al.
Publicado: (2024)
por: Wang, Anna, et al.
Publicado: (2024)
All Neural Low-latency Directional Speech Extraction
por: Pandey, Ashutosh, et al.
Publicado: (2024)
por: Pandey, Ashutosh, et al.
Publicado: (2024)
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
por: Luong, Diep, et al.
Publicado: (2025)
por: Luong, Diep, et al.
Publicado: (2025)
FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2026)
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2026)
MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
por: Shah, Neil, et al.
Publicado: (2024)
por: Shah, Neil, et al.
Publicado: (2024)
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
por: Stahl, Benjamin, et al.
Publicado: (2025)
por: Stahl, Benjamin, et al.
Publicado: (2025)
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
por: Garg, Ashi, et al.
Publicado: (2025)
por: Garg, Ashi, et al.
Publicado: (2025)
Integrated Multi-Level Knowledge Distillation for Enhanced Speaker Verification
por: Yang, Wenhao, et al.
Publicado: (2024)
por: Yang, Wenhao, et al.
Publicado: (2024)
Ejemplares similares
-
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
por: Chi, Hyung Gun, et al.
Publicado: (2025) -
Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
por: Ognjen, et al.
Publicado: (2024) -
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
por: Palaskar, Shruti, et al.
Publicado: (2024) -
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
por: Yuan, Xihao, et al.
Publicado: (2025) -
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
por: Chen, Li-Wei, et al.
Publicado: (2024)