Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lai, Kunfeng, Tang, Zhenheng, Pan, Xinglin, Dong, Peijie, Liu, Xiang, Chen, Haolan, Shen, Li, Li, Bo, Chu, Xiaowen
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915145307914240
author Lai, Kunfeng
Tang, Zhenheng
Pan, Xinglin
Dong, Peijie
Liu, Xiang
Chen, Haolan
Shen, Li
Li, Bo
Chu, Xiaowen
author_facet Lai, Kunfeng
Tang, Zhenheng
Pan, Xinglin
Dong, Peijie
Liu, Xiang
Chen, Haolan
Shen, Li
Li, Bo
Chu, Xiaowen
contents Model merging aggregates Large Language Models (LLMs) finetuned on different tasks into a stronger one. However, parameter conflicts between models leads to performance degradation in averaging. While model routing addresses this issue by selecting individual models during inference, it imposes excessive storage and compute costs, and fails to leverage the common knowledge from different models. In this work, we observe that different layers exhibit varying levels of parameter conflicts. Building on this insight, we average layers with minimal parameter conflicts and use a novel task-level expert routing for layers with significant conflicts. To further reduce storage costs, inspired by task arithmetic sparsity, we decouple multiple fine-tuned experts into a dense expert and several sparse experts. Considering the out-of-distribution samples, we select and merge appropriate experts based on the task uncertainty of the input data. We conduct extensive experiments on both LLaMA and Qwen with varying parameter scales, and evaluate on real-world reasoning tasks. Results demonstrate that our method consistently achieves significant performance improvements while requiring less system cost compared to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04411
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
Lai, Kunfeng
Tang, Zhenheng
Pan, Xinglin
Dong, Peijie
Liu, Xiang
Chen, Haolan
Shen, Li
Li, Bo
Chu, Xiaowen
Machine Learning
Artificial Intelligence
Computation and Language
68T50
Model merging aggregates Large Language Models (LLMs) finetuned on different tasks into a stronger one. However, parameter conflicts between models leads to performance degradation in averaging. While model routing addresses this issue by selecting individual models during inference, it imposes excessive storage and compute costs, and fails to leverage the common knowledge from different models. In this work, we observe that different layers exhibit varying levels of parameter conflicts. Building on this insight, we average layers with minimal parameter conflicts and use a novel task-level expert routing for layers with significant conflicts. To further reduce storage costs, inspired by task arithmetic sparsity, we decouple multiple fine-tuned experts into a dense expert and several sparse experts. Considering the out-of-distribution samples, we select and merge appropriate experts based on the task uncertainty of the input data. We conduct extensive experiments on both LLaMA and Qwen with varying parameter scales, and evaluate on real-world reasoning tasks. Results demonstrate that our method consistently achieves significant performance improvements while requiring less system cost compared to existing methods.
title Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
topic Machine Learning
Artificial Intelligence
Computation and Language
68T50
url https://arxiv.org/abs/2502.04411