AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Yujie, Li, Jian, Dong, Xiaoyu, Xu, Pengfei, Zhou, Xiaohui, Zhang, Yujia, LU, Zexin, Wang, Yasha, Zhao, Alan, Chu, Xu, Wu, Xiao-Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914050456158208
author Feng, Yujie
Li, Jian
Dong, Xiaoyu
Xu, Pengfei
Zhou, Xiaohui
Zhang, Yujia
LU, Zexin
Wang, Yasha
Zhao, Alan
Chu, Xu
Wu, Xiao-Ming
author_facet Feng, Yujie
Li, Jian
Dong, Xiaoyu
Xu, Pengfei
Zhou, Xiaohui
Zhang, Yujia
LU, Zexin
Wang, Yasha
Zhao, Alan
Chu, Xu
Wu, Xiao-Ming
contents Continual learning (CL) is essential for deploying large language models (LLMs) in dynamic real-world environments without the need for costly retraining. Recent model merging-based methods have attracted significant attention, but they still struggle to effectively manage the trade-off between learning new knowledge and preventing forgetting, a challenge largely stemming from suboptimal number of merges and merging frequency. In this paper, we introduce Adaptive Iterative Model Merging (AimMerging), a novel CL framework that utilizes learning and forgetting signals from the training trajectory to dynamically monitor the model's training status. Guided by dynamic monitoring, the training trajectory-guided merge controller adaptively determines the timing and frequency of iterative fusion, while the rehearsal-based knowledge fusion module computes the merging weights and executes the fusion. Comprehensive experiments on three CL benchmarks with various model sizes (from 770M to 13B) demonstrate that AimMerging achieves significant performance improvements over existing state-of-the-art methods, with an average relative improvement of 80% and 59% on FWT and BWT, respectively. The source code is provided for reproducibility.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17348
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning
Feng, Yujie
Li, Jian
Dong, Xiaoyu
Xu, Pengfei
Zhou, Xiaohui
Zhang, Yujia
LU, Zexin
Wang, Yasha
Zhao, Alan
Chu, Xu
Wu, Xiao-Ming
Computation and Language
Artificial Intelligence
Continual learning (CL) is essential for deploying large language models (LLMs) in dynamic real-world environments without the need for costly retraining. Recent model merging-based methods have attracted significant attention, but they still struggle to effectively manage the trade-off between learning new knowledge and preventing forgetting, a challenge largely stemming from suboptimal number of merges and merging frequency. In this paper, we introduce Adaptive Iterative Model Merging (AimMerging), a novel CL framework that utilizes learning and forgetting signals from the training trajectory to dynamically monitor the model's training status. Guided by dynamic monitoring, the training trajectory-guided merge controller adaptively determines the timing and frequency of iterative fusion, while the rehearsal-based knowledge fusion module computes the merging weights and executes the fusion. Comprehensive experiments on three CL benchmarks with various model sizes (from 770M to 13B) demonstrate that AimMerging achieves significant performance improvements over existing state-of-the-art methods, with an average relative improvement of 80% and 59% on FWT and BWT, respectively. The source code is provided for reproducibility.
title AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.17348