Test-Time Model Adaptation for Quantized Neural Networks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Deng, Zeshuai, Chen, Guohao, Niu, Shuaicheng, Luo, Hui, Zhang, Shuhai, Yang, Yifan, Chen, Renjie, Luo, Wei, Tan, Mingkui
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911090350227456
author Deng, Zeshuai
Chen, Guohao
Niu, Shuaicheng
Luo, Hui
Zhang, Shuhai
Yang, Yifan
Chen, Renjie
Luo, Wei
Tan, Mingkui
author_facet Deng, Zeshuai
Chen, Guohao
Niu, Shuaicheng
Luo, Hui
Zhang, Shuhai
Yang, Yifan
Chen, Renjie
Luo, Wei
Tan, Mingkui
contents Quantizing deep models prior to deployment is a widely adopted technique to speed up inference for various real-time applications, such as autonomous driving. However, quantized models often suffer from severe performance degradation in dynamic environments with potential domain shifts and this degradation is significantly more pronounced compared with their full-precision counterparts, as shown by our theoretical and empirical illustrations. To address the domain shift problem, test-time adaptation (TTA) has emerged as an effective solution by enabling models to learn adaptively from test data. Unfortunately, existing TTA methods are often impractical for quantized models as they typically rely on gradient backpropagation--an operation that is unsupported on quantized models due to vanishing gradients, as well as memory and latency constraints. In this paper, we focus on TTA for quantized models to improve their robustness and generalization ability efficiently. We propose a continual zeroth-order adaptation (ZOA) framework that enables efficient model adaptation using only two forward passes, eliminating the computational burden of existing methods. Moreover, we propose a domain knowledge management scheme to store and reuse different domain knowledge with negligible memory consumption, reducing the interference of different domain knowledge and fostering the knowledge accumulation during long-term adaptation. Experimental results on three classical architectures, including quantized transformer-based and CNN-based models, demonstrate the superiority of our methods for quantized model adaptation. On the quantized W6A6 ViT-B model, our ZOA is able to achieve a 5.0\% improvement over the state-of-the-art FOA on ImageNet-C dataset. The source code is available at https://github.com/DengZeshuai/ZOA.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02180
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Test-Time Model Adaptation for Quantized Neural Networks
Deng, Zeshuai
Chen, Guohao
Niu, Shuaicheng
Luo, Hui
Zhang, Shuhai
Yang, Yifan
Chen, Renjie
Luo, Wei
Tan, Mingkui
Computer Vision and Pattern Recognition
Quantizing deep models prior to deployment is a widely adopted technique to speed up inference for various real-time applications, such as autonomous driving. However, quantized models often suffer from severe performance degradation in dynamic environments with potential domain shifts and this degradation is significantly more pronounced compared with their full-precision counterparts, as shown by our theoretical and empirical illustrations. To address the domain shift problem, test-time adaptation (TTA) has emerged as an effective solution by enabling models to learn adaptively from test data. Unfortunately, existing TTA methods are often impractical for quantized models as they typically rely on gradient backpropagation--an operation that is unsupported on quantized models due to vanishing gradients, as well as memory and latency constraints. In this paper, we focus on TTA for quantized models to improve their robustness and generalization ability efficiently. We propose a continual zeroth-order adaptation (ZOA) framework that enables efficient model adaptation using only two forward passes, eliminating the computational burden of existing methods. Moreover, we propose a domain knowledge management scheme to store and reuse different domain knowledge with negligible memory consumption, reducing the interference of different domain knowledge and fostering the knowledge accumulation during long-term adaptation. Experimental results on three classical architectures, including quantized transformer-based and CNN-based models, demonstrate the superiority of our methods for quantized model adaptation. On the quantized W6A6 ViT-B model, our ZOA is able to achieve a 5.0\% improvement over the state-of-the-art FOA on ImageNet-C dataset. The source code is available at https://github.com/DengZeshuai/ZOA.
title Test-Time Model Adaptation for Quantized Neural Networks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.02180