LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Xin, Jia, Hong, Zhou, Hualin, Wang, Sheng Guang, Zhang, Yu, Dang, Ting, Gu, Tao
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910024383594496
author Wang, Xin
Jia, Hong
Zhou, Hualin
Wang, Sheng Guang
Zhang, Yu
Dang, Ting
Gu, Tao
author_facet Wang, Xin
Jia, Hong
Zhou, Hualin
Wang, Sheng Guang
Zhang, Yu
Dang, Ting
Gu, Tao
contents Deploying Vision-Language Models (VLMs) on edge devices is challenged by resource constraints and performance degradation under distribution shifts. While test-time adaptation (TTA) can counteract such shifts, existing methods are too resource-intensive for on-device deployment. To address this challenge, we propose LQA, a lightweight, quantized-adaptive framework for VLMs that combines a modality-aware quantization strategy with gradient-free test-time adaptation. We introduce Selective Hybrid Quantization (SHQ) and a quantized, gradient-free adaptation mechanism to enable robust and efficient VLM deployment on resource-constrained hardware. Experiments across both synthetic and real-world distribution shifts show that LQA improves overall adaptation performance by 4.5\%, uses less memory than full-precision models, and significantly outperforms gradient-based TTA methods, achieving up to 19.9$\times$ lower memory usage across seven open-source datasets. These results demonstrate that LQA offers a practical pathway for robust, privacy-preserving, and efficient VLM deployment on edge devices.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07849
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge
Wang, Xin
Jia, Hong
Zhou, Hualin
Wang, Sheng Guang
Zhang, Yu
Dang, Ting
Gu, Tao
Artificial Intelligence
Deploying Vision-Language Models (VLMs) on edge devices is challenged by resource constraints and performance degradation under distribution shifts. While test-time adaptation (TTA) can counteract such shifts, existing methods are too resource-intensive for on-device deployment. To address this challenge, we propose LQA, a lightweight, quantized-adaptive framework for VLMs that combines a modality-aware quantization strategy with gradient-free test-time adaptation. We introduce Selective Hybrid Quantization (SHQ) and a quantized, gradient-free adaptation mechanism to enable robust and efficient VLM deployment on resource-constrained hardware. Experiments across both synthetic and real-world distribution shifts show that LQA improves overall adaptation performance by 4.5\%, uses less memory than full-precision models, and significantly outperforms gradient-based TTA methods, achieving up to 19.9$\times$ lower memory usage across seven open-source datasets. These results demonstrate that LQA offers a practical pathway for robust, privacy-preserving, and efficient VLM deployment on edge devices.
title LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge
topic Artificial Intelligence
url https://arxiv.org/abs/2602.07849