QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Biswas, Subrata, Khan, Mohammad Nur Hossain, Islam, Bashima
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915447167778816
author Biswas, Subrata
Khan, Mohammad Nur Hossain
Islam, Bashima
author_facet Biswas, Subrata
Khan, Mohammad Nur Hossain
Islam, Bashima
contents Spoken Language Understanding (SLU) systems must balance performance and efficiency, particularly in resource-constrained environments. Existing methods apply distillation and quantization separately, leading to suboptimal compression as distillation ignores quantization constraints. We propose QUADS, a unified framework that optimizes both through multi-stage training with a pre-tuned model, enhancing adaptability to low-bit regimes while maintaining accuracy. QUADS achieves 71.13\% accuracy on SLURP and 99.20\% on FSC, with only minor degradations of up to 5.56\% compared to state-of-the-art models. Additionally, it reduces computational complexity by 60--73$\times$ (GMACs) and model size by 83--700$\times$, demonstrating strong robustness under extreme quantization. These results establish QUADS as a highly efficient solution for real-world, resource-constrained SLU applications.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14723
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
Biswas, Subrata
Khan, Mohammad Nur Hossain
Islam, Bashima
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Sound
Spoken Language Understanding (SLU) systems must balance performance and efficiency, particularly in resource-constrained environments. Existing methods apply distillation and quantization separately, leading to suboptimal compression as distillation ignores quantization constraints. We propose QUADS, a unified framework that optimizes both through multi-stage training with a pre-tuned model, enhancing adaptability to low-bit regimes while maintaining accuracy. QUADS achieves 71.13\% accuracy on SLURP and 99.20\% on FSC, with only minor degradations of up to 5.56\% compared to state-of-the-art models. Additionally, it reduces computational complexity by 60--73$\times$ (GMACs) and model size by 83--700$\times$, demonstrating strong robustness under extreme quantization. These results establish QUADS as a highly efficient solution for real-world, resource-constrained SLU applications.
title QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2505.14723