Task Alignment: A simple and effective proxy for model merging in computer vision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: de Jorge, Pau, de Souza, César Roberto, Michele, Björn, Sarıyıldız, Mert Bülent, Weinzaepfel, Philippe, Perronnin, Florent, Larlus, Diane, Kalantidis, Yannis
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918446370914304
author de Jorge, Pau
de Souza, César Roberto
Michele, Björn
Sarıyıldız, Mert Bülent
Weinzaepfel, Philippe
Perronnin, Florent
Larlus, Diane
Kalantidis, Yannis
author_facet de Jorge, Pau
de Souza, César Roberto
Michele, Björn
Sarıyıldız, Mert Bülent
Weinzaepfel, Philippe
Perronnin, Florent
Larlus, Diane
Kalantidis, Yannis
contents Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Despite extensive prior work, most evaluations of model merging in computer vision are restricted to image classification using CLIP, where different classification datasets define different tasks. In this work, our goal is to make model merging more practical and show its relevance on challenging scenarios beyond this specific setting. In most vision scenarios, different tasks rely on trainable and usually heterogeneous decoders. Differently from previous studies with frozen decoders, where merged models can be evaluated right away, the non-trivial cost of decoder training renders hyperparameter selection based on downstream performance impractical. To address this, we introduce the task alignment proxy, and show how it can be used to speed up hyperparameter selection by orders of magnitude while retaining performance. Equipped with the task alignment proxy, we extend the applicability of model merging to multi-task vision models beyond CLIP-based classification.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12935
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Task Alignment: A simple and effective proxy for model merging in computer vision
de Jorge, Pau
de Souza, César Roberto
Michele, Björn
Sarıyıldız, Mert Bülent
Weinzaepfel, Philippe
Perronnin, Florent
Larlus, Diane
Kalantidis, Yannis
Computer Vision and Pattern Recognition
Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Despite extensive prior work, most evaluations of model merging in computer vision are restricted to image classification using CLIP, where different classification datasets define different tasks. In this work, our goal is to make model merging more practical and show its relevance on challenging scenarios beyond this specific setting. In most vision scenarios, different tasks rely on trainable and usually heterogeneous decoders. Differently from previous studies with frozen decoders, where merged models can be evaluated right away, the non-trivial cost of decoder training renders hyperparameter selection based on downstream performance impractical. To address this, we introduce the task alignment proxy, and show how it can be used to speed up hyperparameter selection by orders of magnitude while retaining performance. Equipped with the task alignment proxy, we extend the applicability of model merging to multi-task vision models beyond CLIP-based classification.
title Task Alignment: A simple and effective proxy for model merging in computer vision
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.12935