Saved in:
Bibliographic Details
Main Authors: Yang, Zhaorui, Pang, Tianyu, Feng, Haozhe, Wang, Han, Chen, Wei, Zhu, Minfeng, Liu, Qian
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2402.13669
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917677347373056
author Yang, Zhaorui
Pang, Tianyu
Feng, Haozhe
Wang, Han
Chen, Wei
Zhu, Minfeng
Liu, Qian
author_facet Yang, Zhaorui
Pang, Tianyu
Feng, Haozhe
Wang, Han
Chen, Wei
Zhu, Minfeng
Liu, Qian
contents The surge in Large Language Models (LLMs) has revolutionized natural language processing, but fine-tuning them for specific tasks often encounters challenges in balancing performance and preserving general instruction-following abilities. In this paper, we posit that the distribution gap between task datasets and the LLMs serves as the primary underlying cause. To address the problem, we introduce Self-Distillation Fine-Tuning (SDFT), a novel approach that bridges the distribution gap by guiding fine-tuning with a distilled dataset generated by the model itself to match its original distribution. Experimental results on the Llama-2-chat model across various benchmarks demonstrate that SDFT effectively mitigates catastrophic forgetting while achieving comparable or superior performance on downstream tasks compared to the vanilla fine-tuning. Moreover, SDFT demonstrates the potential to maintain the helpfulness and safety alignment of LLMs. Our code is available at https://github.com/sail-sg/sdft.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13669
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning
Yang, Zhaorui
Pang, Tianyu
Feng, Haozhe
Wang, Han
Chen, Wei
Zhu, Minfeng
Liu, Qian
Computation and Language
The surge in Large Language Models (LLMs) has revolutionized natural language processing, but fine-tuning them for specific tasks often encounters challenges in balancing performance and preserving general instruction-following abilities. In this paper, we posit that the distribution gap between task datasets and the LLMs serves as the primary underlying cause. To address the problem, we introduce Self-Distillation Fine-Tuning (SDFT), a novel approach that bridges the distribution gap by guiding fine-tuning with a distilled dataset generated by the model itself to match its original distribution. Experimental results on the Llama-2-chat model across various benchmarks demonstrate that SDFT effectively mitigates catastrophic forgetting while achieving comparable or superior performance on downstream tasks compared to the vanilla fine-tuning. Moreover, SDFT demonstrates the potential to maintain the helpfulness and safety alignment of LLMs. Our code is available at https://github.com/sail-sg/sdft.
title Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning
topic Computation and Language
url https://arxiv.org/abs/2402.13669