Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yaze, Liu, Yicong, Zou, Yixiong, Li, Yuhua, Li, Ruixuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916004023500800
author Zhao, Yaze
Liu, Yicong
Zou, Yixiong
Li, Yuhua
Li, Ruixuan
author_facet Zhao, Yaze
Liu, Yicong
Zou, Yixiong
Li, Yuhua
Li, Ruixuan
contents Cross-Domain Few-Shot Learning (CDFSL) aims to adapt large-scale pretrained models to specialized target domains with limited samples, yet the few-shot fine-tuning of vision-language models like CLIP remains underexplored. By establishing multiple fine-tuning baselines of CLIP for CDFSL, we find adapter-based methods (e.g., LoRA) consistently outperform prompt-based ones (e.g., MaPLe), contrary to in-domain scenarios. To make those effective in-domain methods competitive again in CDFSL, we analyze this phenomenon and discover LoRA's superiority stems from rectifying the collapsed attention of visual CLS token, enhancing modality alignment and class separation by focusing on text-related visual regions. Further, we find textual EOS token exhibit much better attention to visual samples, and CLIP's standard contrastive loss weakly constrains modality alignment. Based on these insights, we propose Semantic Probe, a plug-and-play attention rectification framework for both adapter- and prompt-based methods. Extensive experiments on four CDFSL benchmarks validate our rationale, achieving state-of-the-art performance and benefiting both fine-tuning paradigms. Codes will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11659
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning
Zhao, Yaze
Liu, Yicong
Zou, Yixiong
Li, Yuhua
Li, Ruixuan
Computer Vision and Pattern Recognition
Artificial Intelligence
Cross-Domain Few-Shot Learning (CDFSL) aims to adapt large-scale pretrained models to specialized target domains with limited samples, yet the few-shot fine-tuning of vision-language models like CLIP remains underexplored. By establishing multiple fine-tuning baselines of CLIP for CDFSL, we find adapter-based methods (e.g., LoRA) consistently outperform prompt-based ones (e.g., MaPLe), contrary to in-domain scenarios. To make those effective in-domain methods competitive again in CDFSL, we analyze this phenomenon and discover LoRA's superiority stems from rectifying the collapsed attention of visual CLS token, enhancing modality alignment and class separation by focusing on text-related visual regions. Further, we find textual EOS token exhibit much better attention to visual samples, and CLIP's standard contrastive loss weakly constrains modality alignment. Based on these insights, we propose Semantic Probe, a plug-and-play attention rectification framework for both adapter- and prompt-based methods. Extensive experiments on four CDFSL benchmarks validate our rationale, achieving state-of-the-art performance and benefiting both fine-tuning paradigms. Codes will be released.
title Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.11659