ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Junseok, Kim, Nahun, Lee, Sangyong, Chun, Chang-Jae |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
di: Lee, Junseok, et al.
Pubblicazione: (2026)
di: Lee, Junseok, et al.
Pubblicazione: (2026)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2026)
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2026)
Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
di: Shendabadi, Ali, et al.
Pubblicazione: (2026)
di: Shendabadi, Ali, et al.
Pubblicazione: (2026)
Raon-Speech Technical Report
di: Kim, Beomsoo, et al.
Pubblicazione: (2026)
di: Kim, Beomsoo, et al.
Pubblicazione: (2026)
Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
di: Wang, Chien-Chun, et al.
Pubblicazione: (2024)
di: Wang, Chien-Chun, et al.
Pubblicazione: (2024)
Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages
di: Pillai, Leena G, et al.
Pubblicazione: (2024)
di: Pillai, Leena G, et al.
Pubblicazione: (2024)
Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt Tuning
di: Yang, Hongli, et al.
Pubblicazione: (2025)
di: Yang, Hongli, et al.
Pubblicazione: (2025)
WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition
di: Ramezani, Erfan, et al.
Pubblicazione: (2026)
di: Ramezani, Erfan, et al.
Pubblicazione: (2026)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
di: Shao, Hang, et al.
Pubblicazione: (2023)
di: Shao, Hang, et al.
Pubblicazione: (2023)
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring
di: Sudarshan, Ankitha, et al.
Pubblicazione: (2023)
di: Sudarshan, Ankitha, et al.
Pubblicazione: (2023)
Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
di: Radhakrishnan, Srijith, et al.
Pubblicazione: (2023)
di: Radhakrishnan, Srijith, et al.
Pubblicazione: (2023)
Benchmarking Automatic Speech Recognition for Indian Languages in Agricultural Contexts
di: S, Chandrashekar M, et al.
Pubblicazione: (2026)
di: S, Chandrashekar M, et al.
Pubblicazione: (2026)
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition
di: Su, Hsuan, et al.
Pubblicazione: (2024)
di: Su, Hsuan, et al.
Pubblicazione: (2024)
EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs
di: Lin, Liang, et al.
Pubblicazione: (2026)
di: Lin, Liang, et al.
Pubblicazione: (2026)
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
di: Nassereldine, Amir, et al.
Pubblicazione: (2024)
di: Nassereldine, Amir, et al.
Pubblicazione: (2024)
Efficient Streaming LLM for Speech Recognition
di: Jia, Junteng, et al.
Pubblicazione: (2024)
di: Jia, Junteng, et al.
Pubblicazione: (2024)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
di: Yang, Chenchen, et al.
Pubblicazione: (2026)
di: Yang, Chenchen, et al.
Pubblicazione: (2026)
Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition
di: Gheffari, Youcef Soufiane, et al.
Pubblicazione: (2026)
di: Gheffari, Youcef Soufiane, et al.
Pubblicazione: (2026)
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
di: Wang, Xiong, et al.
Pubblicazione: (2024)
di: Wang, Xiong, et al.
Pubblicazione: (2024)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
di: He, Xinlu, et al.
Pubblicazione: (2025)
di: He, Xinlu, et al.
Pubblicazione: (2025)
Efficient Dialect-Aware Modeling and Conditioning for Low-Resource Taiwanese Hakka Speech Processing
di: Peng, An-Ci, et al.
Pubblicazione: (2026)
di: Peng, An-Ci, et al.
Pubblicazione: (2026)
Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
di: Zhang, Wei, et al.
Pubblicazione: (2025)
di: Zhang, Wei, et al.
Pubblicazione: (2025)
Contextual Earnings-22: A Speech Recognition Benchmark with Custom Vocabulary in the Wild
di: Durmus, Berkin, et al.
Pubblicazione: (2026)
di: Durmus, Berkin, et al.
Pubblicazione: (2026)
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
di: Sy, Yaya, et al.
Pubblicazione: (2025)
di: Sy, Yaya, et al.
Pubblicazione: (2025)
Efficient Training for Cross-lingual Speech Language Models
di: Zhou, Yan, et al.
Pubblicazione: (2026)
di: Zhou, Yan, et al.
Pubblicazione: (2026)
An Investigation Into Explainable Audio Hate Speech Detection
di: An, Jinmyeong, et al.
Pubblicazione: (2024)
di: An, Jinmyeong, et al.
Pubblicazione: (2024)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
di: Wang, Peng, et al.
Pubblicazione: (2026)
di: Wang, Peng, et al.
Pubblicazione: (2026)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
di: Jo, Daejin, et al.
Pubblicazione: (2025)
di: Jo, Daejin, et al.
Pubblicazione: (2025)
Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
di: Lee, Dongwook, et al.
Pubblicazione: (2026)
di: Lee, Dongwook, et al.
Pubblicazione: (2026)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
di: Wang, Yujin, et al.
Pubblicazione: (2022)
di: Wang, Yujin, et al.
Pubblicazione: (2022)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
di: Ginjala, Srishti, et al.
Pubblicazione: (2026)
di: Ginjala, Srishti, et al.
Pubblicazione: (2026)
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
di: Zhao, Qiuming, et al.
Pubblicazione: (2025)
di: Zhao, Qiuming, et al.
Pubblicazione: (2025)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2024)
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2024)
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024)
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024)
Length-Aware Rotary Position Embedding for Text-Speech Alignment
di: Kim, Hyeongju, et al.
Pubblicazione: (2025)
di: Kim, Hyeongju, et al.
Pubblicazione: (2025)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
di: Rufai, Amina Mardiyyah, et al.
Pubblicazione: (2020)
di: Rufai, Amina Mardiyyah, et al.
Pubblicazione: (2020)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
di: Sheng, Zhengyan, et al.
Pubblicazione: (2025)
di: Sheng, Zhengyan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
di: Lee, Junseok, et al.
Pubblicazione: (2026) -
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2026) -
Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
di: Shendabadi, Ali, et al.
Pubblicazione: (2026) -
Raon-Speech Technical Report
di: Kim, Beomsoo, et al.
Pubblicazione: (2026) -
Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
di: Wang, Chien-Chun, et al.
Pubblicazione: (2024)