GeneralizeFormer: Layer-Adaptive Model Generation across Test-Time Distribution Shifts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ambekar, Sameer, Xiao, Zehao, Zhen, Xiantong, Snoek, Cees G. M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909499263025152
author Ambekar, Sameer
Xiao, Zehao
Zhen, Xiantong
Snoek, Cees G. M.
author_facet Ambekar, Sameer
Xiao, Zehao
Zhen, Xiantong
Snoek, Cees G. M.
contents We consider the problem of test-time domain generalization, where a model is trained on several source domains and adjusted on target domains never seen during training. Different from the common methods that fine-tune the model or adjust the classifier parameters online, we propose to generate multiple layer parameters on the fly during inference by a lightweight meta-learned transformer, which we call \textit{GeneralizeFormer}. The layer-wise parameters are generated per target batch without fine-tuning or online adjustment. By doing so, our method is more effective in dynamic scenarios with multiple target distributions and also avoids forgetting valuable source distribution characteristics. Moreover, by considering layer-wise gradients, the proposed method adapts itself to various distribution shifts. To reduce the computational and time cost, we fix the convolutional parameters while only generating parameters of the Batch Normalization layers and the linear classifier. Experiments on six widely used domain generalization datasets demonstrate the benefits and abilities of the proposed method to efficiently handle various distribution shifts, generalize in dynamic scenarios, and avoid forgetting.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12195
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeneralizeFormer: Layer-Adaptive Model Generation across Test-Time Distribution Shifts
Ambekar, Sameer
Xiao, Zehao
Zhen, Xiantong
Snoek, Cees G. M.
Machine Learning
We consider the problem of test-time domain generalization, where a model is trained on several source domains and adjusted on target domains never seen during training. Different from the common methods that fine-tune the model or adjust the classifier parameters online, we propose to generate multiple layer parameters on the fly during inference by a lightweight meta-learned transformer, which we call \textit{GeneralizeFormer}. The layer-wise parameters are generated per target batch without fine-tuning or online adjustment. By doing so, our method is more effective in dynamic scenarios with multiple target distributions and also avoids forgetting valuable source distribution characteristics. Moreover, by considering layer-wise gradients, the proposed method adapts itself to various distribution shifts. To reduce the computational and time cost, we fix the convolutional parameters while only generating parameters of the Batch Normalization layers and the linear classifier. Experiments on six widely used domain generalization datasets demonstrate the benefits and abilities of the proposed method to efficiently handle various distribution shifts, generalize in dynamic scenarios, and avoid forgetting.
title GeneralizeFormer: Layer-Adaptive Model Generation across Test-Time Distribution Shifts
topic Machine Learning
url https://arxiv.org/abs/2502.12195