Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Suzuki, Satoshi, Yamaguchi, Shin'ya, Takeda, Shoichiro, Yamane, Taiga, Makishima, Naoki, Kawata, Naotaka, Ihori, Mana, Tanaka, Tomohiro, Orihashi, Shota, Masumura, Ryo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917078180560896
author Suzuki, Satoshi
Yamaguchi, Shin'ya
Takeda, Shoichiro
Yamane, Taiga
Makishima, Naoki
Kawata, Naotaka
Ihori, Mana
Tanaka, Tomohiro
Orihashi, Shota
Masumura, Ryo
author_facet Suzuki, Satoshi
Yamaguchi, Shin'ya
Takeda, Shoichiro
Yamane, Taiga
Makishima, Naoki
Kawata, Naotaka
Ihori, Mana
Tanaka, Tomohiro
Orihashi, Shota
Masumura, Ryo
contents Contrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image and text encoders. This paper aims to robustly fine-tune these vision-language models on in-distribution (ID) data without compromising their generalization abilities in out-of-distribution (OOD) and zero-shot settings. Current robust fine-tuning methods tackle this challenge by reusing contrastive learning, which was used in pre-training, for fine-tuning. However, we found that these methods distort the geometric structure of the embeddings, which plays a crucial role in the generalization of vision-language models, resulting in limited OOD and zero-shot performance. To address this, we propose Difference Vector Equalization (DiVE), which preserves the geometric structure during fine-tuning. The idea behind DiVE is to constrain difference vectors, each of which is obtained by subtracting the embeddings extracted from the pre-trained and fine-tuning models for the same data sample. By constraining the difference vectors to be equal across various data samples, we effectively preserve the geometric structure. Therefore, we introduce two losses: average vector loss (AVL) and pairwise vector loss (PVL). AVL preserves the geometric structure globally by constraining difference vectors to be equal to their weighted average. PVL preserves the geometric structure locally by ensuring a consistent multimodal alignment. Our experiments demonstrate that DiVE effectively preserves the geometric structure, achieving strong results across ID, OOD, and zero-shot metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09973
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models
Suzuki, Satoshi
Yamaguchi, Shin'ya
Takeda, Shoichiro
Yamane, Taiga
Makishima, Naoki
Kawata, Naotaka
Ihori, Mana
Tanaka, Tomohiro
Orihashi, Shota
Masumura, Ryo
Computer Vision and Pattern Recognition
Artificial Intelligence
Contrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image and text encoders. This paper aims to robustly fine-tune these vision-language models on in-distribution (ID) data without compromising their generalization abilities in out-of-distribution (OOD) and zero-shot settings. Current robust fine-tuning methods tackle this challenge by reusing contrastive learning, which was used in pre-training, for fine-tuning. However, we found that these methods distort the geometric structure of the embeddings, which plays a crucial role in the generalization of vision-language models, resulting in limited OOD and zero-shot performance. To address this, we propose Difference Vector Equalization (DiVE), which preserves the geometric structure during fine-tuning. The idea behind DiVE is to constrain difference vectors, each of which is obtained by subtracting the embeddings extracted from the pre-trained and fine-tuning models for the same data sample. By constraining the difference vectors to be equal across various data samples, we effectively preserve the geometric structure. Therefore, we introduce two losses: average vector loss (AVL) and pairwise vector loss (PVL). AVL preserves the geometric structure globally by constraining difference vectors to be equal to their weighted average. PVL preserves the geometric structure locally by ensuring a consistent multimodal alignment. Our experiments demonstrate that DiVE effectively preserves the geometric structure, achieving strong results across ID, OOD, and zero-shot metrics.
title Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.09973