StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Younghyun, Jeong, Jongheon, Kwak, Sangkyung, Lee, Kyungmin, Lee, Juho, Shin, Jinwoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909661350854656
author Kim, Younghyun
Jeong, Jongheon
Kwak, Sangkyung
Lee, Kyungmin
Lee, Juho
Shin, Jinwoo
author_facet Kim, Younghyun
Jeong, Jongheon
Kwak, Sangkyung
Lee, Kyungmin
Lee, Juho
Shin, Jinwoo
contents Learning robust representations from data often requires scale, which has led to the success of recent zero-shot models such as CLIP. However, the obtained robustness can easily be deteriorated when these models are fine-tuned on other downstream tasks (e.g., of smaller scales). Previous works often interpret this phenomenon in the context of domain shift, developing fine-tuning methods that aim to preserve the original domain as much as possible. However, in a different context, fine-tuned models with limited data are also prone to learning features that are spurious to humans, such as background or texture. In this paper, we propose StarFT (Spurious Textual Alignment Regularization), a novel framework for fine-tuning zero-shot models to enhance robustness by preventing them from learning spuriosity. We introduce a regularization that aligns the output distribution for spuriosity-injected labels with the original zero-shot model, ensuring that the model is not induced to extract irrelevant features further from these descriptions. We leverage recent language models to get such spuriosity-injected labels by generating alternative textual descriptions that highlight potentially confounding features. Extensive experiments validate the robust generalization of StarFT and its emerging properties: zero-shot group robustness and improved zero-shot classification. Notably, StarFT boosts both worst-group and average accuracy by 14.30% and 3.02%, respectively, in the Waterbirds group shift scenario, where other robust fine-tuning baselines show even degraded performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13232
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment
Kim, Younghyun
Jeong, Jongheon
Kwak, Sangkyung
Lee, Kyungmin
Lee, Juho
Shin, Jinwoo
Artificial Intelligence
Computer Vision and Pattern Recognition
Learning robust representations from data often requires scale, which has led to the success of recent zero-shot models such as CLIP. However, the obtained robustness can easily be deteriorated when these models are fine-tuned on other downstream tasks (e.g., of smaller scales). Previous works often interpret this phenomenon in the context of domain shift, developing fine-tuning methods that aim to preserve the original domain as much as possible. However, in a different context, fine-tuned models with limited data are also prone to learning features that are spurious to humans, such as background or texture. In this paper, we propose StarFT (Spurious Textual Alignment Regularization), a novel framework for fine-tuning zero-shot models to enhance robustness by preventing them from learning spuriosity. We introduce a regularization that aligns the output distribution for spuriosity-injected labels with the original zero-shot model, ensuring that the model is not induced to extract irrelevant features further from these descriptions. We leverage recent language models to get such spuriosity-injected labels by generating alternative textual descriptions that highlight potentially confounding features. Extensive experiments validate the robust generalization of StarFT and its emerging properties: zero-shot group robustness and improved zero-shot classification. Notably, StarFT boosts both worst-group and average accuracy by 14.30% and 3.02%, respectively, in the Waterbirds group shift scenario, where other robust fine-tuning baselines show even degraded performance.
title StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.13232