Exploiting Layer Normalization Fine-tuning in Visual Transformer Foundation Models for Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Zhaorui, Pan, Tan, Huang, Kaizhu, Yu, Weimiao, Yao, Kai, Jiang, Chen, Wang, Qiufeng, Nguyen, Anh, Guo, Xin, Cheng, Yuan, Yang, Xi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911100819210240
author Tan, Zhaorui
Pan, Tan
Huang, Kaizhu
Yu, Weimiao
Yao, Kai
Jiang, Chen
Wang, Qiufeng
Nguyen, Anh
Guo, Xin
Cheng, Yuan
Yang, Xi
author_facet Tan, Zhaorui
Pan, Tan
Huang, Kaizhu
Yu, Weimiao
Yao, Kai
Jiang, Chen
Wang, Qiufeng
Nguyen, Anh
Guo, Xin
Cheng, Yuan
Yang, Xi
contents LayerNorm is pivotal in Vision Transformers (ViTs), yet its fine-tuning dynamics under data scarcity and domain shifts remain underexplored. This paper shows that shifts in LayerNorm parameters after fine-tuning (LayerNorm shifts) are indicative of the transitions between source and target domains; its efficacy is contingent upon the degree to which the target training samples accurately represent the target domain, as quantified by our proposed Fine-tuning Shift Ratio ($FSR$). Building on this, we propose a simple yet effective rescaling mechanism using a scalar $λ$ that is negatively correlated to $FSR$ to align learned LayerNorm shifts with those ideal shifts achieved under fully representative data, combined with a cyclic framework that further enhances the LayerNorm fine-tuning. Extensive experiments across natural and pathological images, in both in-distribution (ID) and out-of-distribution (OOD) settings, and various target training sample regimes validate our framework. Notably, OOD tasks tend to yield lower $FSR$ and higher $λ$ in comparison to ID cases, especially with scarce data, indicating under-represented target training samples. Moreover, ViTFs fine-tuned on pathological data behave more like ID settings, favoring conservative LayerNorm updates. Our findings illuminate the underexplored dynamics of LayerNorm in transfer learning and provide practical strategies for LayerNorm fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07577
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploiting Layer Normalization Fine-tuning in Visual Transformer Foundation Models for Classification
Tan, Zhaorui
Pan, Tan
Huang, Kaizhu
Yu, Weimiao
Yao, Kai
Jiang, Chen
Wang, Qiufeng
Nguyen, Anh
Guo, Xin
Cheng, Yuan
Yang, Xi
Computer Vision and Pattern Recognition
Machine Learning
LayerNorm is pivotal in Vision Transformers (ViTs), yet its fine-tuning dynamics under data scarcity and domain shifts remain underexplored. This paper shows that shifts in LayerNorm parameters after fine-tuning (LayerNorm shifts) are indicative of the transitions between source and target domains; its efficacy is contingent upon the degree to which the target training samples accurately represent the target domain, as quantified by our proposed Fine-tuning Shift Ratio ($FSR$). Building on this, we propose a simple yet effective rescaling mechanism using a scalar $λ$ that is negatively correlated to $FSR$ to align learned LayerNorm shifts with those ideal shifts achieved under fully representative data, combined with a cyclic framework that further enhances the LayerNorm fine-tuning. Extensive experiments across natural and pathological images, in both in-distribution (ID) and out-of-distribution (OOD) settings, and various target training sample regimes validate our framework. Notably, OOD tasks tend to yield lower $FSR$ and higher $λ$ in comparison to ID cases, especially with scarce data, indicating under-represented target training samples. Moreover, ViTFs fine-tuned on pathological data behave more like ID settings, favoring conservative LayerNorm updates. Our findings illuminate the underexplored dynamics of LayerNorm in transfer learning and provide practical strategies for LayerNorm fine-tuning.
title Exploiting Layer Normalization Fine-tuning in Visual Transformer Foundation Models for Classification
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2508.07577