LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910024383594496 |
|---|---|
| author | Wang, Xin Jia, Hong Zhou, Hualin Wang, Sheng Guang Zhang, Yu Dang, Ting Gu, Tao |
| author_facet | Wang, Xin Jia, Hong Zhou, Hualin Wang, Sheng Guang Zhang, Yu Dang, Ting Gu, Tao |
| contents | Deploying Vision-Language Models (VLMs) on edge devices is challenged by resource constraints and performance degradation under distribution shifts. While test-time adaptation (TTA) can counteract such shifts, existing methods are too resource-intensive for on-device deployment. To address this challenge, we propose LQA, a lightweight, quantized-adaptive framework for VLMs that combines a modality-aware quantization strategy with gradient-free test-time adaptation. We introduce Selective Hybrid Quantization (SHQ) and a quantized, gradient-free adaptation mechanism to enable robust and efficient VLM deployment on resource-constrained hardware. Experiments across both synthetic and real-world distribution shifts show that LQA improves overall adaptation performance by 4.5\%, uses less memory than full-precision models, and significantly outperforms gradient-based TTA methods, achieving up to 19.9$\times$ lower memory usage across seven open-source datasets. These results demonstrate that LQA offers a practical pathway for robust, privacy-preserving, and efficient VLM deployment on edge devices. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_07849 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge Wang, Xin Jia, Hong Zhou, Hualin Wang, Sheng Guang Zhang, Yu Dang, Ting Gu, Tao Artificial Intelligence Deploying Vision-Language Models (VLMs) on edge devices is challenged by resource constraints and performance degradation under distribution shifts. While test-time adaptation (TTA) can counteract such shifts, existing methods are too resource-intensive for on-device deployment. To address this challenge, we propose LQA, a lightweight, quantized-adaptive framework for VLMs that combines a modality-aware quantization strategy with gradient-free test-time adaptation. We introduce Selective Hybrid Quantization (SHQ) and a quantized, gradient-free adaptation mechanism to enable robust and efficient VLM deployment on resource-constrained hardware. Experiments across both synthetic and real-world distribution shifts show that LQA improves overall adaptation performance by 4.5\%, uses less memory than full-precision models, and significantly outperforms gradient-based TTA methods, achieving up to 19.9$\times$ lower memory usage across seven open-source datasets. These results demonstrate that LQA offers a practical pathway for robust, privacy-preserving, and efficient VLM deployment on edge devices. |
| title | LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2602.07849 |