QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915447167778816 |
|---|---|
| author | Biswas, Subrata Khan, Mohammad Nur Hossain Islam, Bashima |
| author_facet | Biswas, Subrata Khan, Mohammad Nur Hossain Islam, Bashima |
| contents | Spoken Language Understanding (SLU) systems must balance performance and efficiency, particularly in resource-constrained environments. Existing methods apply distillation and quantization separately, leading to suboptimal compression as distillation ignores quantization constraints. We propose QUADS, a unified framework that optimizes both through multi-stage training with a pre-tuned model, enhancing adaptability to low-bit regimes while maintaining accuracy. QUADS achieves 71.13\% accuracy on SLURP and 99.20\% on FSC, with only minor degradations of up to 5.56\% compared to state-of-the-art models. Additionally, it reduces computational complexity by 60--73$\times$ (GMACs) and model size by 83--700$\times$, demonstrating strong robustness under extreme quantization. These results establish QUADS as a highly efficient solution for real-world, resource-constrained SLU applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14723 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding Biswas, Subrata Khan, Mohammad Nur Hossain Islam, Bashima Audio and Speech Processing Artificial Intelligence Computation and Language Machine Learning Sound Spoken Language Understanding (SLU) systems must balance performance and efficiency, particularly in resource-constrained environments. Existing methods apply distillation and quantization separately, leading to suboptimal compression as distillation ignores quantization constraints. We propose QUADS, a unified framework that optimizes both through multi-stage training with a pre-tuned model, enhancing adaptability to low-bit regimes while maintaining accuracy. QUADS achieves 71.13\% accuracy on SLURP and 99.20\% on FSC, with only minor degradations of up to 5.56\% compared to state-of-the-art models. Additionally, it reduces computational complexity by 60--73$\times$ (GMACs) and model size by 83--700$\times$, demonstrating strong robustness under extreme quantization. These results establish QUADS as a highly efficient solution for real-world, resource-constrained SLU applications. |
| title | QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding |
| topic | Audio and Speech Processing Artificial Intelligence Computation and Language Machine Learning Sound |
| url | https://arxiv.org/abs/2505.14723 |