Salvato in:
Dettagli Bibliografici
Autori principali: Fish, Edward, Michieli, Umberto, Ozay, Mete
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:https://arxiv.org/abs/2307.12659
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929239841832960
author Fish, Edward
Michieli, Umberto
Ozay, Mete
author_facet Fish, Edward
Michieli, Umberto
Ozay, Mete
contents Recent advancement in Automatic Speech Recognition (ASR) has produced large AI models, which become impractical for deployment in mobile devices. Model quantization is effective to produce compressed general-purpose models, however such models may only be deployed to a restricted sub-domain of interest. We show that ASR models can be personalized during quantization while relying on just a small set of unlabelled samples from the target domain. To this end, we propose myQASR, a mixed-precision quantization method that generates tailored quantization schemes for diverse users under any memory requirement with no fine-tuning. myQASR automatically evaluates the quantization sensitivity of network layers by analysing the full-precision activation values. We are then able to generate a personalised mixed-precision quantization scheme for any pre-determined memory budget. Results for large-scale ASR models show how myQASR improves performance for specific genders, languages, and speakers.
format Preprint
id arxiv_https___arxiv_org_abs_2307_12659
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Model for Every User and Budget: Label-Free and Personalized Mixed-Precision Quantization
Fish, Edward
Michieli, Umberto
Ozay, Mete
Sound
Computation and Language
Audio and Speech Processing
Recent advancement in Automatic Speech Recognition (ASR) has produced large AI models, which become impractical for deployment in mobile devices. Model quantization is effective to produce compressed general-purpose models, however such models may only be deployed to a restricted sub-domain of interest. We show that ASR models can be personalized during quantization while relying on just a small set of unlabelled samples from the target domain. To this end, we propose myQASR, a mixed-precision quantization method that generates tailored quantization schemes for diverse users under any memory requirement with no fine-tuning. myQASR automatically evaluates the quantization sensitivity of network layers by analysing the full-precision activation values. We are then able to generate a personalised mixed-precision quantization scheme for any pre-determined memory budget. Results for large-scale ASR models show how myQASR improves performance for specific genders, languages, and speakers.
title A Model for Every User and Budget: Label-Free and Personalized Mixed-Precision Quantization
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2307.12659