Model Merging by Uncertainty-Based Gradient Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Daheim, Nico, Möllenhoff, Thomas, Ponti, Edoardo Maria, Gurevych, Iryna, Khan, Mohammad Emtiyaz
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914921088811008
author Daheim, Nico
Möllenhoff, Thomas
Ponti, Edoardo Maria
Gurevych, Iryna
Khan, Mohammad Emtiyaz
author_facet Daheim, Nico
Möllenhoff, Thomas
Ponti, Edoardo Maria
Gurevych, Iryna
Khan, Mohammad Emtiyaz
contents Models trained on different datasets can be merged by a weighted-averaging of their parameters, but why does it work and when can it fail? Here, we connect the inaccuracy of weighted-averaging to mismatches in the gradients and propose a new uncertainty-based scheme to improve the performance by reducing the mismatch. The connection also reveals implicit assumptions in other schemes such as averaging, task arithmetic, and Fisher-weighted averaging. Our new method gives consistent improvements for large language models and vision transformers, both in terms of performance and robustness to hyperparameters. Code available here.
format Preprint
id arxiv_https___arxiv_org_abs_2310_12808
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Model Merging by Uncertainty-Based Gradient Matching
Daheim, Nico
Möllenhoff, Thomas
Ponti, Edoardo Maria
Gurevych, Iryna
Khan, Mohammad Emtiyaz
Machine Learning
Artificial Intelligence
Computation and Language
Models trained on different datasets can be merged by a weighted-averaging of their parameters, but why does it work and when can it fail? Here, we connect the inaccuracy of weighted-averaging to mismatches in the gradients and propose a new uncertainty-based scheme to improve the performance by reducing the mismatch. The connection also reveals implicit assumptions in other schemes such as averaging, task arithmetic, and Fisher-weighted averaging. Our new method gives consistent improvements for large language models and vision transformers, both in terms of performance and robustness to hyperparameters. Code available here.
title Model Merging by Uncertainty-Based Gradient Matching
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2310.12808