MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909853592584192 |
|---|---|
| author | Chen, Yupeng Wang, Senmiao Zhang, Yushun Lin, Zhihang Zhang, Haozhe Sun, Weijian Ding, Tian Sun, Ruoyu |
| author_facet | Chen, Yupeng Wang, Senmiao Zhang, Yushun Lin, Zhihang Zhang, Haozhe Sun, Weijian Ding, Tian Sun, Ruoyu |
| contents | Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. Typically, LLMs are first pre-trained on large corpora and subsequently fine-tuned on task-specific datasets. However, during fine-tuning, LLMs may forget some knowledge acquired in the pre-training stage, leading to a decline in general capabilities. Existing approaches to mitigate forgetting often rely on access to pre-training data, which may be unavailable in many real-world scenarios--such as fine-tuning checkpoint-only open-source LLMs. To address this challenge, we propose a new fine-tuning algorithm termed Momentum-Filtered Optimizer (MoFO). MoFO is an extension of greedy block coordinate descent (BCD) methods: in each iteration, MoFO only updates the model parameters with the largest momentum magnitudes, while keeping all other parameters fixed. MoFO achieves similar fine-tuning performance to the default fine-tuning algorithm while effectively mitigating knowledge forgetting. We validate MoFO through rigorous convergence analysis and extensive experiments, demonstrating its effectiveness in mitigating forgetting without pre-training data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_20999 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning Chen, Yupeng Wang, Senmiao Zhang, Yushun Lin, Zhihang Zhang, Haozhe Sun, Weijian Ding, Tian Sun, Ruoyu Machine Learning Artificial Intelligence Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. Typically, LLMs are first pre-trained on large corpora and subsequently fine-tuned on task-specific datasets. However, during fine-tuning, LLMs may forget some knowledge acquired in the pre-training stage, leading to a decline in general capabilities. Existing approaches to mitigate forgetting often rely on access to pre-training data, which may be unavailable in many real-world scenarios--such as fine-tuning checkpoint-only open-source LLMs. To address this challenge, we propose a new fine-tuning algorithm termed Momentum-Filtered Optimizer (MoFO). MoFO is an extension of greedy block coordinate descent (BCD) methods: in each iteration, MoFO only updates the model parameters with the largest momentum magnitudes, while keeping all other parameters fixed. MoFO achieves similar fine-tuning performance to the default fine-tuning algorithm while effectively mitigating knowledge forgetting. We validate MoFO through rigorous convergence analysis and extensive experiments, demonstrating its effectiveness in mitigating forgetting without pre-training data. |
| title | MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2407.20999 |