Dual Knowledge Distillation for Efficient Sound Event Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Yang, Das, Rohan Kumar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UCIL: An Unsupervised Class Incremental Learning Approach for Sound Event Detection
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
von: Biswas, Subrata, et al.
Veröffentlicht: (2025)
von: Biswas, Subrata, et al.
Veröffentlicht: (2025)
WildDESED: An LLM-Powered Dataset for Wild Domestic Environment Sound Event Detection System
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
von: Zhang, Shiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Shiqi, et al.
Veröffentlicht: (2025)
EnvSDD: Benchmarking Environmental Sound Deepfake Detection
von: Yin, Han, et al.
Veröffentlicht: (2025)
von: Yin, Han, et al.
Veröffentlicht: (2025)
SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
von: Hussain, Aafiya, et al.
Veröffentlicht: (2026)
von: Hussain, Aafiya, et al.
Veröffentlicht: (2026)
Modality-Invariant Bidirectional Temporal Representation Distillation Network for Missing Multimodal Sentiment Analysis
von: Wang, Xincheng, et al.
Veröffentlicht: (2025)
von: Wang, Xincheng, et al.
Veröffentlicht: (2025)
Leveraging LLM and Text-Queried Separation for Noise-Robust Sound Event Detection
von: Yin, Han, et al.
Veröffentlicht: (2024)
von: Yin, Han, et al.
Veröffentlicht: (2024)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Multimodal Sentiment Analysis with Missing Modality: A Knowledge-Transfer Approach
von: Liu, Weide, et al.
Veröffentlicht: (2023)
von: Liu, Weide, et al.
Veröffentlicht: (2023)
Imagine to Hear: Auditory Knowledge Generation can be an Effective Assistant for Language Models
von: Yoo, Suho, et al.
Veröffentlicht: (2025)
von: Yoo, Suho, et al.
Veröffentlicht: (2025)
Label-Looping: Highly Efficient Decoding for Transducers
von: Bataev, Vladimir, et al.
Veröffentlicht: (2024)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2024)
FlashSpeech: Efficient Zero-Shot Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
On the Role of Speech Data in Reducing Toxicity Detection Bias
von: Bell, Samuel J., et al.
Veröffentlicht: (2024)
von: Bell, Samuel J., et al.
Veröffentlicht: (2024)
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
von: Yang, Yiyuan, et al.
Veröffentlicht: (2025)
von: Yang, Yiyuan, et al.
Veröffentlicht: (2025)
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
von: Inoue, Nakamasa, et al.
Veröffentlicht: (2024)
von: Inoue, Nakamasa, et al.
Veröffentlicht: (2024)
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
von: Wang, Qiaolin, et al.
Veröffentlicht: (2025)
von: Wang, Qiaolin, et al.
Veröffentlicht: (2025)
SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models
von: Kumar, Anurag, et al.
Veröffentlicht: (2025)
von: Kumar, Anurag, et al.
Veröffentlicht: (2025)
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
von: Abdelfattah, Abdullah, et al.
Veröffentlicht: (2025)
von: Abdelfattah, Abdullah, et al.
Veröffentlicht: (2025)
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection
von: Fujita, Yoto, et al.
Veröffentlicht: (2024)
von: Fujita, Yoto, et al.
Veröffentlicht: (2024)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
FMSG-JLESS Submission for DCASE 2024 Task4 on Sound Event Detection with Heterogeneous Training Dataset and Potentially Missing Labels
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet
von: Wang, Shenran, et al.
Veröffentlicht: (2025)
von: Wang, Shenran, et al.
Veröffentlicht: (2025)
Where's That Voice Coming? Continual Learning for Sound Source Localization
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2023)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2023)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
TF-Mamba: A Time-Frequency Network for Sound Source Localization
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2025)
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2025)
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
von: Ju, Zeqian, et al.
Veröffentlicht: (2024)
von: Ju, Zeqian, et al.
Veröffentlicht: (2024)
MoonCast: High-Quality Zero-Shot Podcast Generation
von: Ju, Zeqian, et al.
Veröffentlicht: (2025)
von: Ju, Zeqian, et al.
Veröffentlicht: (2025)
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
von: Li, Yizhi, et al.
Veröffentlicht: (2023)
von: Li, Yizhi, et al.
Veröffentlicht: (2023)
DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
von: Yan, Sheng, et al.
Veröffentlicht: (2024)
von: Yan, Sheng, et al.
Veröffentlicht: (2024)
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
von: Weise, Tobias, et al.
Veröffentlicht: (2024)
von: Weise, Tobias, et al.
Veröffentlicht: (2024)
MIMII-Agent: Leveraging LLMs with Function Calling for Relative Evaluation of Anomalous Sound Detection
von: Purohit, Harsh, et al.
Veröffentlicht: (2025)
von: Purohit, Harsh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UCIL: An Unsupervised Class Incremental Learning Approach for Sound Event Detection
von: Xiao, Yang, et al.
Veröffentlicht: (2024) -
QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
von: Biswas, Subrata, et al.
Veröffentlicht: (2025) -
WildDESED: An LLM-Powered Dataset for Wild Domestic Environment Sound Event Detection System
von: Xiao, Yang, et al.
Veröffentlicht: (2024) -
Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
von: Zhang, Shiqi, et al.
Veröffentlicht: (2025) -
EnvSDD: Benchmarking Environmental Sound Deepfake Detection
von: Yin, Han, et al.
Veröffentlicht: (2025)