Recurrent Knowledge Identification and Fusion for Language Model Continual Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Yujie, Wang, Xujia, Lu, Zexin, Fu, Shenghong, Shi, Guangyuan, Xu, Yongxin, Wang, Yasha, Yu, Philip S., Chu, Xu, Wu, Xiao-Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912403187302400
author Feng, Yujie
Wang, Xujia
Lu, Zexin
Fu, Shenghong
Shi, Guangyuan
Xu, Yongxin
Wang, Yasha
Yu, Philip S.
Chu, Xu
Wu, Xiao-Ming
author_facet Feng, Yujie
Wang, Xujia
Lu, Zexin
Fu, Shenghong
Shi, Guangyuan
Xu, Yongxin
Wang, Yasha
Yu, Philip S.
Chu, Xu
Wu, Xiao-Ming
contents Continual learning (CL) is crucial for deploying large language models (LLMs) in dynamic real-world environments without costly retraining. While recent model ensemble and model merging methods guided by parameter importance have gained popularity, they often struggle to balance knowledge transfer and forgetting, mainly due to the reliance on static importance estimates during sequential training. In this paper, we present Recurrent-KIF, a novel CL framework for Recurrent Knowledge Identification and Fusion, which enables dynamic estimation of parameter importance distributions to enhance knowledge transfer. Inspired by human continual learning, Recurrent-KIF employs an inner loop that rapidly adapts to new tasks while identifying important parameters, coupled with an outer loop that globally manages the fusion of new and historical knowledge through redundant knowledge pruning and key knowledge merging. These inner-outer loops iteratively perform multiple rounds of fusion, allowing Recurrent-KIF to leverage intermediate training information and adaptively adjust fusion strategies based on evolving importance distributions. Extensive experiments on two CL benchmarks with various model sizes (from 770M to 13B) demonstrate that Recurrent-KIF effectively mitigates catastrophic forgetting and enhances knowledge transfer.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17510
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Recurrent Knowledge Identification and Fusion for Language Model Continual Learning
Feng, Yujie
Wang, Xujia
Lu, Zexin
Fu, Shenghong
Shi, Guangyuan
Xu, Yongxin
Wang, Yasha
Yu, Philip S.
Chu, Xu
Wu, Xiao-Ming
Machine Learning
Artificial Intelligence
Computation and Language
Continual learning (CL) is crucial for deploying large language models (LLMs) in dynamic real-world environments without costly retraining. While recent model ensemble and model merging methods guided by parameter importance have gained popularity, they often struggle to balance knowledge transfer and forgetting, mainly due to the reliance on static importance estimates during sequential training. In this paper, we present Recurrent-KIF, a novel CL framework for Recurrent Knowledge Identification and Fusion, which enables dynamic estimation of parameter importance distributions to enhance knowledge transfer. Inspired by human continual learning, Recurrent-KIF employs an inner loop that rapidly adapts to new tasks while identifying important parameters, coupled with an outer loop that globally manages the fusion of new and historical knowledge through redundant knowledge pruning and key knowledge merging. These inner-outer loops iteratively perform multiple rounds of fusion, allowing Recurrent-KIF to leverage intermediate training information and adaptively adjust fusion strategies based on evolving importance distributions. Extensive experiments on two CL benchmarks with various model sizes (from 770M to 13B) demonstrate that Recurrent-KIF effectively mitigates catastrophic forgetting and enhances knowledge transfer.
title Recurrent Knowledge Identification and Fusion for Language Model Continual Learning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.17510