MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Yupeng, Wang, Senmiao, Zhang, Yushun, Lin, Zhihang, Zhang, Haozhe, Sun, Weijian, Ding, Tian, Sun, Ruoyu
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909853592584192
author Chen, Yupeng
Wang, Senmiao
Zhang, Yushun
Lin, Zhihang
Zhang, Haozhe
Sun, Weijian
Ding, Tian
Sun, Ruoyu
author_facet Chen, Yupeng
Wang, Senmiao
Zhang, Yushun
Lin, Zhihang
Zhang, Haozhe
Sun, Weijian
Ding, Tian
Sun, Ruoyu
contents Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. Typically, LLMs are first pre-trained on large corpora and subsequently fine-tuned on task-specific datasets. However, during fine-tuning, LLMs may forget some knowledge acquired in the pre-training stage, leading to a decline in general capabilities. Existing approaches to mitigate forgetting often rely on access to pre-training data, which may be unavailable in many real-world scenarios--such as fine-tuning checkpoint-only open-source LLMs. To address this challenge, we propose a new fine-tuning algorithm termed Momentum-Filtered Optimizer (MoFO). MoFO is an extension of greedy block coordinate descent (BCD) methods: in each iteration, MoFO only updates the model parameters with the largest momentum magnitudes, while keeping all other parameters fixed. MoFO achieves similar fine-tuning performance to the default fine-tuning algorithm while effectively mitigating knowledge forgetting. We validate MoFO through rigorous convergence analysis and extensive experiments, demonstrating its effectiveness in mitigating forgetting without pre-training data.
format Preprint
id arxiv_https___arxiv_org_abs_2407_20999
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning
Chen, Yupeng
Wang, Senmiao
Zhang, Yushun
Lin, Zhihang
Zhang, Haozhe
Sun, Weijian
Ding, Tian
Sun, Ruoyu
Machine Learning
Artificial Intelligence
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. Typically, LLMs are first pre-trained on large corpora and subsequently fine-tuned on task-specific datasets. However, during fine-tuning, LLMs may forget some knowledge acquired in the pre-training stage, leading to a decline in general capabilities. Existing approaches to mitigate forgetting often rely on access to pre-training data, which may be unavailable in many real-world scenarios--such as fine-tuning checkpoint-only open-source LLMs. To address this challenge, we propose a new fine-tuning algorithm termed Momentum-Filtered Optimizer (MoFO). MoFO is an extension of greedy block coordinate descent (BCD) methods: in each iteration, MoFO only updates the model parameters with the largest momentum magnitudes, while keeping all other parameters fixed. MoFO achieves similar fine-tuning performance to the default fine-tuning algorithm while effectively mitigating knowledge forgetting. We validate MoFO through rigorous convergence analysis and extensive experiments, demonstrating its effectiveness in mitigating forgetting without pre-training data.
title MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2407.20999