DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shao, Hang, Liu, Bei, Wang, Wei, Gong, Xun, Qian, Yanmin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023)
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023)
Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
von: Liu, Bei, et al.
Veröffentlicht: (2024)
von: Liu, Bei, et al.
Veröffentlicht: (2024)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation
von: Le, Chenyang, et al.
Veröffentlicht: (2025)
von: Le, Chenyang, et al.
Veröffentlicht: (2025)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
von: Gong, Xun, et al.
Veröffentlicht: (2024)
von: Gong, Xun, et al.
Veröffentlicht: (2024)
Efficient Multilingual ASR Finetuning via LoRA Language Experts
von: Li, Jiahong, et al.
Veröffentlicht: (2025)
von: Li, Jiahong, et al.
Veröffentlicht: (2025)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
von: Gong, Xun, et al.
Veröffentlicht: (2025)
von: Gong, Xun, et al.
Veröffentlicht: (2025)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
von: Bijoy, Mehedi Hasan, et al.
Veröffentlicht: (2025)
von: Bijoy, Mehedi Hasan, et al.
Veröffentlicht: (2025)
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
von: Wei, Kun, et al.
Veröffentlicht: (2023)
von: Wei, Kun, et al.
Veröffentlicht: (2023)
MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition
von: Xia, Yinfeng, et al.
Veröffentlicht: (2025)
von: Xia, Yinfeng, et al.
Veröffentlicht: (2025)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
Sparsely Shared LoRA on Whisper for Child Speech Recognition
von: Liu, Wei, et al.
Veröffentlicht: (2023)
von: Liu, Wei, et al.
Veröffentlicht: (2023)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
von: Li, Jinpeng, et al.
Veröffentlicht: (2024)
von: Li, Jinpeng, et al.
Veröffentlicht: (2024)
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
von: Nassereldine, Amir, et al.
Veröffentlicht: (2024)
von: Nassereldine, Amir, et al.
Veröffentlicht: (2024)
Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker Verification
von: Liu, Bei, et al.
Veröffentlicht: (2024)
von: Liu, Bei, et al.
Veröffentlicht: (2024)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
von: Huang, Wen, et al.
Veröffentlicht: (2025)
von: Huang, Wen, et al.
Veröffentlicht: (2025)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Prompting Whisper for Joint Speech Transcription and Diarization
von: Zamyrova, Mariia, et al.
Veröffentlicht: (2026)
von: Zamyrova, Mariia, et al.
Veröffentlicht: (2026)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus
von: Shekoufandeh, Golshid, et al.
Veröffentlicht: (2025)
von: Shekoufandeh, Golshid, et al.
Veröffentlicht: (2025)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
von: Zhuo, Le, et al.
Veröffentlicht: (2023)
von: Zhuo, Le, et al.
Veröffentlicht: (2023)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
Adapting Where It Matters: Depth-Aware Adaptation for Efficient Multilingual Speech Recognition in Low-Resource Languages
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Text adaptation for speaker verification with speaker-text factorized embeddings
von: Yang, Yexin, et al.
Veröffentlicht: (2025)
von: Yang, Yexin, et al.
Veröffentlicht: (2025)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
Quantizing Whisper-small: How design choices affect ASR performance
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025)
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
von: Shao, Qijie, et al.
Veröffentlicht: (2025) -
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023) -
Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
von: Liu, Bei, et al.
Veröffentlicht: (2024) -
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
von: Meng, Lingwei, et al.
Veröffentlicht: (2024) -
Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation
von: Le, Chenyang, et al.
Veröffentlicht: (2025)