MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Lu, Zhang, Tianyu, Bu, Zhiqi, Wang, Suyuchen, He, Huan, Fu, Jie, Wu, Yonghui, Bian, Jiang, Chen, Yong, Bengio, Yoshua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912345447464960
author Li, Lu
Zhang, Tianyu
Bu, Zhiqi
Wang, Suyuchen
He, Huan
Fu, Jie
Wu, Yonghui
Bian, Jiang
Chen, Yong
Bengio, Yoshua
author_facet Li, Lu
Zhang, Tianyu
Bu, Zhiqi
Wang, Suyuchen
He, Huan
Fu, Jie
Wu, Yonghui
Bian, Jiang
Chen, Yong
Bengio, Yoshua
contents Model merging has emerged as an effective approach to combine multiple single-task models into a multitask model. This process typically involves computing a weighted average of the model parameters without any additional training. Existing model-merging methods focus on enhancing average task accuracy. However, interference and conflicts between the objectives of different tasks can lead to trade-offs during the merging process. In real-world applications, a set of solutions with various trade-offs can be more informative, helping practitioners make decisions based on diverse preferences. In this paper, we introduce a novel and low-compute algorithm, Model Merging with Amortized Pareto Front (MAP). MAP efficiently identifies a Pareto set of scaling coefficients for merging multiple models, reflecting the trade-offs involved. It amortizes the substantial computational cost of evaluations needed to estimate the Pareto front by using quadratic approximation surrogate models derived from a pre-selected set of scaling coefficients. Experimental results on vision and natural language processing tasks demonstrate that MAP can accurately identify the Pareto front, providing practitioners with flexible solutions to balance competing task objectives. We also introduce Bayesian MAP for scenarios with a relatively low number of tasks and Nested MAP for situations with a high number of tasks, further reducing the computational cost of evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2406_07529
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation
Li, Lu
Zhang, Tianyu
Bu, Zhiqi
Wang, Suyuchen
He, Huan
Fu, Jie
Wu, Yonghui
Bian, Jiang
Chen, Yong
Bengio, Yoshua
Machine Learning
Model merging has emerged as an effective approach to combine multiple single-task models into a multitask model. This process typically involves computing a weighted average of the model parameters without any additional training. Existing model-merging methods focus on enhancing average task accuracy. However, interference and conflicts between the objectives of different tasks can lead to trade-offs during the merging process. In real-world applications, a set of solutions with various trade-offs can be more informative, helping practitioners make decisions based on diverse preferences. In this paper, we introduce a novel and low-compute algorithm, Model Merging with Amortized Pareto Front (MAP). MAP efficiently identifies a Pareto set of scaling coefficients for merging multiple models, reflecting the trade-offs involved. It amortizes the substantial computational cost of evaluations needed to estimate the Pareto front by using quadratic approximation surrogate models derived from a pre-selected set of scaling coefficients. Experimental results on vision and natural language processing tasks demonstrate that MAP can accurately identify the Pareto front, providing practitioners with flexible solutions to balance competing task objectives. We also introduce Bayesian MAP for scenarios with a relatively low number of tasks and Nested MAP for situations with a high number of tasks, further reducing the computational cost of evaluation.
title MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation
topic Machine Learning
url https://arxiv.org/abs/2406.07529