Training-free LLM Merging for Multi-task Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fu, Zichuan, Wu, Xian, Wang, Yejing, Wang, Wanyu, Ye, Shanshan, Yin, Hongzhi, Chang, Yi, Zheng, Yefeng, Zhao, Xiangyu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911006484070400
author Fu, Zichuan
Wu, Xian
Wang, Yejing
Wang, Wanyu
Ye, Shanshan
Yin, Hongzhi
Chang, Yi
Zheng, Yefeng
Zhao, Xiangyu
author_facet Fu, Zichuan
Wu, Xian
Wang, Yejing
Wang, Wanyu
Ye, Shanshan
Yin, Hongzhi
Chang, Yi
Zheng, Yefeng
Zhao, Xiangyu
contents Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse natural language processing (NLP) tasks. The release of open-source LLMs like LLaMA and Qwen has triggered the development of numerous fine-tuned models tailored for various tasks and languages. In this paper, we explore an important question: is it possible to combine these specialized models to create a unified model with multi-task capabilities. We introduces Hierarchical Iterative Merging (Hi-Merging), a training-free method for unifying different specialized LLMs into a single model. Specifically, Hi-Merging employs model-wise and layer-wise pruning and scaling, guided by contribution analysis, to mitigate parameter conflicts. Extensive experiments on multiple-choice and question-answering tasks in both Chinese and English validate Hi-Merging's ability for multi-task learning. The results demonstrate that Hi-Merging consistently outperforms existing merging techniques and surpasses the performance of models fine-tuned on combined datasets in most scenarios. Code is available at: https://github.com/Applied-Machine-Learning-Lab/Hi-Merging.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training-free LLM Merging for Multi-task Learning
Fu, Zichuan
Wu, Xian
Wang, Yejing
Wang, Wanyu
Ye, Shanshan
Yin, Hongzhi
Chang, Yi
Zheng, Yefeng
Zhao, Xiangyu
Computation and Language
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse natural language processing (NLP) tasks. The release of open-source LLMs like LLaMA and Qwen has triggered the development of numerous fine-tuned models tailored for various tasks and languages. In this paper, we explore an important question: is it possible to combine these specialized models to create a unified model with multi-task capabilities. We introduces Hierarchical Iterative Merging (Hi-Merging), a training-free method for unifying different specialized LLMs into a single model. Specifically, Hi-Merging employs model-wise and layer-wise pruning and scaling, guided by contribution analysis, to mitigate parameter conflicts. Extensive experiments on multiple-choice and question-answering tasks in both Chinese and English validate Hi-Merging's ability for multi-task learning. The results demonstrate that Hi-Merging consistently outperforms existing merging techniques and surpasses the performance of models fine-tuned on combined datasets in most scenarios. Code is available at: https://github.com/Applied-Machine-Learning-Lab/Hi-Merging.
title Training-free LLM Merging for Multi-task Learning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.12379