Quantizing Whisper-small: How design choices affect ASR performance

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Söhler, Arthur, Irigoyen, Julian, Kirkedal, Andreas Søeborg
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913150522097664
author Söhler, Arthur
Irigoyen, Julian
Kirkedal, Andreas Søeborg
author_facet Söhler, Arthur
Irigoyen, Julian
Kirkedal, Andreas Søeborg
contents Large speech recognition models like Whisper-small achieve high accuracy but are difficult to deploy on edge devices due to their high computational demand. To this end, we present a unified, cross-library evaluation of post-training quantization (PTQ) on Whisper-small that disentangles the impact of quantization scheme, method, granularity, and bit-width. Our study is based on four libraries: PyTorch, Optimum-Quanto, HQQ, and bitsandbytes. Experiments on LibriSpeech test-clean and test-other show that dynamic int8 quantization with Quanto offers the best trade-off, reducing model size by 57% while improving on the baseline's word error rate. Static quantization performed worse, likely due to Whisper's Transformer architecture, while more aggressive formats (e.g., nf4, int3) achieved up to 71% compression at the cost of accuracy in noisy conditions. Overall, our results demonstrate that carefully chosen PTQ methods can substantially reduce model size and inference cost without retraining, enabling efficient deployment of Whisper-small on constrained hardware.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08093
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quantizing Whisper-small: How design choices affect ASR performance
Söhler, Arthur
Irigoyen, Julian
Kirkedal, Andreas Søeborg
Audio and Speech Processing
Computation and Language
Sound
Large speech recognition models like Whisper-small achieve high accuracy but are difficult to deploy on edge devices due to their high computational demand. To this end, we present a unified, cross-library evaluation of post-training quantization (PTQ) on Whisper-small that disentangles the impact of quantization scheme, method, granularity, and bit-width. Our study is based on four libraries: PyTorch, Optimum-Quanto, HQQ, and bitsandbytes. Experiments on LibriSpeech test-clean and test-other show that dynamic int8 quantization with Quanto offers the best trade-off, reducing model size by 57% while improving on the baseline's word error rate. Static quantization performed worse, likely due to Whisper's Transformer architecture, while more aggressive formats (e.g., nf4, int3) achieved up to 71% compression at the cost of accuracy in noisy conditions. Overall, our results demonstrate that carefully chosen PTQ methods can substantially reduce model size and inference cost without retraining, enabling efficient deployment of Whisper-small on constrained hardware.
title Quantizing Whisper-small: How design choices affect ASR performance
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2511.08093