DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910623831425024 |
|---|---|
| author | Shao, Hang Liu, Bei Wang, Wei Gong, Xun Qian, Yanmin |
| author_facet | Shao, Hang Liu, Bei Wang, Wei Gong, Xun Qian, Yanmin |
| contents | As a popular multilingual and multitask pre-trained speech model, Whisper has the problem of curse of multilinguality. To enhance multilingual capabilities in small Whisper models, we propose DQ-Whisper, a novel joint distillation and quantization framework to compress Whisper for efficient inference. Firstly, we propose a novel dynamic matching distillation strategy. Then, a quantization-aware distillation framework is introduced to integrate quantization with distillation. Experimental results on various multilingual datasets show that our suggested distillation approach can effectively enhance the multilingual capabilities of small Whisper models without increasing computational costs. Up to 5.18x reduction in model size is achieved with marginal performance degradation. In addition, quantization is compatible with distillation, which can result in a higher compression rate. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_10788 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition Shao, Hang Liu, Bei Wang, Wei Gong, Xun Qian, Yanmin Sound Computation and Language Audio and Speech Processing As a popular multilingual and multitask pre-trained speech model, Whisper has the problem of curse of multilinguality. To enhance multilingual capabilities in small Whisper models, we propose DQ-Whisper, a novel joint distillation and quantization framework to compress Whisper for efficient inference. Firstly, we propose a novel dynamic matching distillation strategy. Then, a quantization-aware distillation framework is introduced to integrate quantization with distillation. Experimental results on various multilingual datasets show that our suggested distillation approach can effectively enhance the multilingual capabilities of small Whisper models without increasing computational costs. Up to 5.18x reduction in model size is achieved with marginal performance degradation. In addition, quantization is compatible with distillation, which can result in a higher compression rate. |
| title | DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2305.10788 |