How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ming, Yifei, Li, Yixuan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913449472163840
author Ming, Yifei
Li, Yixuan
author_facet Ming, Yifei
Li, Yixuan
contents Recent large vision-language models such as CLIP have shown remarkable out-of-distribution (OOD) detection and generalization performance. However, their zero-shot in-distribution (ID) accuracy is often limited for downstream datasets. Recent CLIP-based fine-tuning methods such as prompt learning have demonstrated significant improvements in ID classification and OOD generalization where OOD labels are available. Nonetheless, it remains unclear whether the model is reliable to semantic shifts without OOD labels. In this paper, we aim to bridge the gap and present a comprehensive study to understand how fine-tuning impact OOD detection for few-shot downstream tasks. By framing OOD detection as multi-modal concept matching, we establish a connection between fine-tuning methods and various OOD scores. Our results suggest that a proper choice of OOD scores is essential for CLIP-based fine-tuning. In particular, the maximum concept matching (MCM) score provides a promising solution consistently. We also show that prompt learning demonstrates the state-of-the-art OOD detection performance over the zero-shot counterpart.
format Preprint
id arxiv_https___arxiv_org_abs_2306_06048
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?
Ming, Yifei
Li, Yixuan
Computer Vision and Pattern Recognition
Computers and Society
Machine Learning
Recent large vision-language models such as CLIP have shown remarkable out-of-distribution (OOD) detection and generalization performance. However, their zero-shot in-distribution (ID) accuracy is often limited for downstream datasets. Recent CLIP-based fine-tuning methods such as prompt learning have demonstrated significant improvements in ID classification and OOD generalization where OOD labels are available. Nonetheless, it remains unclear whether the model is reliable to semantic shifts without OOD labels. In this paper, we aim to bridge the gap and present a comprehensive study to understand how fine-tuning impact OOD detection for few-shot downstream tasks. By framing OOD detection as multi-modal concept matching, we establish a connection between fine-tuning methods and various OOD scores. Our results suggest that a proper choice of OOD scores is essential for CLIP-based fine-tuning. In particular, the maximum concept matching (MCM) score provides a promising solution consistently. We also show that prompt learning demonstrates the state-of-the-art OOD detection performance over the zero-shot counterpart.
title How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?
topic Computer Vision and Pattern Recognition
Computers and Society
Machine Learning
url https://arxiv.org/abs/2306.06048