FeatCal: Feature Calibration for Post-Merging Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Yanggan, Cai, Shuo, Wang, Zihao, Wang, Wenjun, Wang, Yuanyi, Wang, Pengkai, Huang, Sirui, Lu, Su, Wu, Jianmin, Yang, Hongxia
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913122413969408
author Gu, Yanggan
Cai, Shuo
Wang, Zihao
Wang, Wenjun
Wang, Yuanyi
Wang, Pengkai
Huang, Sirui
Lu, Su
Wu, Jianmin
Yang, Hongxia
author_facet Gu, Yanggan
Cai, Shuo
Wang, Zihao
Wang, Wenjun
Wang, Yuanyi
Wang, Pengkai
Huang, Sirui
Lu, Su
Wu, Jianmin
Yang, Hongxia
contents Model merging combines task experts into one model and avoids joint training, retraining, or deploying many expert models, but the merged model often still underperforms task experts. We study this performance gap through feature drift, the difference between features produced by the merged model and by the expert on the same input. Our theory decomposes this drift into upstream propagation and local mismatch, tracks how it propagates and combines through later layers in forward order, and links final feature drift to output drift. This view motivates FeatCal, which uses a small calibration set to calibrate the merged model weights layer by layer in forward order, reducing feature drift while staying close to merged weights and preserving the benefits of model merging. FeatCal uses an efficient closed-form solution to update model weights, with no gradient descent, iterative optimization, or extra modules. On the main CLIP and GLUE benchmarks, FeatCal beats Surgery and ProbSurgery, the closest post-merging calibration baselines: 85.5% vs. 77.0%/78.8% on CLIP-ViT-B/32 Task Arithmetic (TA) and 85.2% vs. 83.7%/82.2% on FLAN-T5-base GLUE. On CLIP-ViT-B/32, 8 examples per task reach 82.9%, and 256 examples per task take 53 seconds, about 4x faster than both baselines, showing better sample efficiency and lower calibration cost.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13030
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FeatCal: Feature Calibration for Post-Merging Models
Gu, Yanggan
Cai, Shuo
Wang, Zihao
Wang, Wenjun
Wang, Yuanyi
Wang, Pengkai
Huang, Sirui
Lu, Su
Wu, Jianmin
Yang, Hongxia
Machine Learning
Artificial Intelligence
Model merging combines task experts into one model and avoids joint training, retraining, or deploying many expert models, but the merged model often still underperforms task experts. We study this performance gap through feature drift, the difference between features produced by the merged model and by the expert on the same input. Our theory decomposes this drift into upstream propagation and local mismatch, tracks how it propagates and combines through later layers in forward order, and links final feature drift to output drift. This view motivates FeatCal, which uses a small calibration set to calibrate the merged model weights layer by layer in forward order, reducing feature drift while staying close to merged weights and preserving the benefits of model merging. FeatCal uses an efficient closed-form solution to update model weights, with no gradient descent, iterative optimization, or extra modules. On the main CLIP and GLUE benchmarks, FeatCal beats Surgery and ProbSurgery, the closest post-merging calibration baselines: 85.5% vs. 77.0%/78.8% on CLIP-ViT-B/32 Task Arithmetic (TA) and 85.2% vs. 83.7%/82.2% on FLAN-T5-base GLUE. On CLIP-ViT-B/32, 8 examples per task reach 82.9%, and 256 examples per task take 53 seconds, about 4x faster than both baselines, showing better sample efficiency and lower calibration cost.
title FeatCal: Feature Calibration for Post-Merging Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.13030