BaFTA: Backprop-Free Test-Time Adaptation For Zero-Shot Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Xuefeng, Zhang, Ke, Sun, Min, Chen, Albert, Kuo, Cheng-Hao, Nevatia, Ram
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916290271117312
author Hu, Xuefeng
Zhang, Ke
Sun, Min
Chen, Albert
Kuo, Cheng-Hao
Nevatia, Ram
author_facet Hu, Xuefeng
Zhang, Ke
Sun, Min
Chen, Albert
Kuo, Cheng-Hao
Nevatia, Ram
contents Large-scale pretrained vision-language models like CLIP have demonstrated remarkable zero-shot image classification capabilities across diverse domains. To enhance CLIP's performance while preserving the zero-shot paradigm, various test-time prompt tuning methods have been introduced to refine class embeddings through unsupervised learning objectives during inference. However, these methods often encounter challenges in selecting appropriate learning rates to prevent collapsed training in the absence of validation data during test-time adaptation. In this study, we propose a novel backpropagation-free algorithm BaFTA for test-time adaptation of vision-language models. Instead of fine-tuning text prompts to refine class embeddings, our approach directly estimates class centroids using online clustering within a projected embedding space that aligns text and visual embeddings. We dynamically aggregate predictions from both estimated and original class embeddings, as well as from distinct augmented views, by assessing the reliability of each prediction using Rényi Entropy. Through extensive experiments, we demonstrate that BaFTA consistently outperforms state-of-the-art test-time adaptation methods in both effectiveness and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11309
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BaFTA: Backprop-Free Test-Time Adaptation For Zero-Shot Vision-Language Models
Hu, Xuefeng
Zhang, Ke
Sun, Min
Chen, Albert
Kuo, Cheng-Hao
Nevatia, Ram
Computer Vision and Pattern Recognition
Large-scale pretrained vision-language models like CLIP have demonstrated remarkable zero-shot image classification capabilities across diverse domains. To enhance CLIP's performance while preserving the zero-shot paradigm, various test-time prompt tuning methods have been introduced to refine class embeddings through unsupervised learning objectives during inference. However, these methods often encounter challenges in selecting appropriate learning rates to prevent collapsed training in the absence of validation data during test-time adaptation. In this study, we propose a novel backpropagation-free algorithm BaFTA for test-time adaptation of vision-language models. Instead of fine-tuning text prompts to refine class embeddings, our approach directly estimates class centroids using online clustering within a projected embedding space that aligns text and visual embeddings. We dynamically aggregate predictions from both estimated and original class embeddings, as well as from distinct augmented views, by assessing the reliability of each prediction using Rényi Entropy. Through extensive experiments, we demonstrate that BaFTA consistently outperforms state-of-the-art test-time adaptation methods in both effectiveness and efficiency.
title BaFTA: Backprop-Free Test-Time Adaptation For Zero-Shot Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.11309