DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shao, Hang, Liu, Bei, Wang, Wei, Gong, Xun, Qian, Yanmin
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910623831425024
author Shao, Hang
Liu, Bei
Wang, Wei
Gong, Xun
Qian, Yanmin
author_facet Shao, Hang
Liu, Bei
Wang, Wei
Gong, Xun
Qian, Yanmin
contents As a popular multilingual and multitask pre-trained speech model, Whisper has the problem of curse of multilinguality. To enhance multilingual capabilities in small Whisper models, we propose DQ-Whisper, a novel joint distillation and quantization framework to compress Whisper for efficient inference. Firstly, we propose a novel dynamic matching distillation strategy. Then, a quantization-aware distillation framework is introduced to integrate quantization with distillation. Experimental results on various multilingual datasets show that our suggested distillation approach can effectively enhance the multilingual capabilities of small Whisper models without increasing computational costs. Up to 5.18x reduction in model size is achieved with marginal performance degradation. In addition, quantization is compatible with distillation, which can result in a higher compression rate.
format Preprint
id arxiv_https___arxiv_org_abs_2305_10788
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
Shao, Hang
Liu, Bei
Wang, Wei
Gong, Xun
Qian, Yanmin
Sound
Computation and Language
Audio and Speech Processing
As a popular multilingual and multitask pre-trained speech model, Whisper has the problem of curse of multilinguality. To enhance multilingual capabilities in small Whisper models, we propose DQ-Whisper, a novel joint distillation and quantization framework to compress Whisper for efficient inference. Firstly, we propose a novel dynamic matching distillation strategy. Then, a quantization-aware distillation framework is introduced to integrate quantization with distillation. Experimental results on various multilingual datasets show that our suggested distillation approach can effectively enhance the multilingual capabilities of small Whisper models without increasing computational costs. Up to 5.18x reduction in model size is achieved with marginal performance degradation. In addition, quantization is compatible with distillation, which can result in a higher compression rate.
title DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2305.10788