Towards Calibrated Robust Fine-Tuning of Vision-Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Oh, Changdae, Lim, Hyesu, Kim, Mijoo, Han, Dongyoon, Yun, Sangdoo, Choo, Jaegul, Hauptmann, Alexander, Cheng, Zhi-Qi, Song, Kyungwoo
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915007949701120
author Oh, Changdae
Lim, Hyesu
Kim, Mijoo
Han, Dongyoon
Yun, Sangdoo
Choo, Jaegul
Hauptmann, Alexander
Cheng, Zhi-Qi
Song, Kyungwoo
author_facet Oh, Changdae
Lim, Hyesu
Kim, Mijoo
Han, Dongyoon
Yun, Sangdoo
Choo, Jaegul
Hauptmann, Alexander
Cheng, Zhi-Qi
Song, Kyungwoo
contents Improving out-of-distribution (OOD) generalization during in-distribution (ID) adaptation is a primary goal of robust fine-tuning of zero-shot models beyond naive fine-tuning. However, despite decent OOD generalization performance from recent robust fine-tuning methods, confidence calibration for reliable model output has not been fully addressed. This work proposes a robust fine-tuning method that improves both OOD accuracy and confidence calibration simultaneously in vision language models. Firstly, we show that both OOD classification and OOD calibration errors have a shared upper bound consisting of two terms of ID data: 1) ID calibration error and 2) the smallest singular value of the ID input covariance matrix. Based on this insight, we design a novel framework that conducts fine-tuning with a constrained multimodal contrastive loss enforcing a larger smallest singular value, which is further guided by the self-distillation of a moving-averaged model to achieve calibrated prediction as well. Starting from empirical evidence supporting our theoretical statements, we provide extensive experimental results on ImageNet distribution shift benchmarks that demonstrate the effectiveness of our theorem and its practical implementation.
format Preprint
id arxiv_https___arxiv_org_abs_2311_01723
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards Calibrated Robust Fine-Tuning of Vision-Language Models
Oh, Changdae
Lim, Hyesu
Kim, Mijoo
Han, Dongyoon
Yun, Sangdoo
Choo, Jaegul
Hauptmann, Alexander
Cheng, Zhi-Qi
Song, Kyungwoo
Computer Vision and Pattern Recognition
Artificial Intelligence
Improving out-of-distribution (OOD) generalization during in-distribution (ID) adaptation is a primary goal of robust fine-tuning of zero-shot models beyond naive fine-tuning. However, despite decent OOD generalization performance from recent robust fine-tuning methods, confidence calibration for reliable model output has not been fully addressed. This work proposes a robust fine-tuning method that improves both OOD accuracy and confidence calibration simultaneously in vision language models. Firstly, we show that both OOD classification and OOD calibration errors have a shared upper bound consisting of two terms of ID data: 1) ID calibration error and 2) the smallest singular value of the ID input covariance matrix. Based on this insight, we design a novel framework that conducts fine-tuning with a constrained multimodal contrastive loss enforcing a larger smallest singular value, which is further guided by the self-distillation of a moving-averaged model to achieve calibrated prediction as well. Starting from empirical evidence supporting our theoretical statements, we provide extensive experimental results on ImageNet distribution shift benchmarks that demonstrate the effectiveness of our theorem and its practical implementation.
title Towards Calibrated Robust Fine-Tuning of Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2311.01723