SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Yuxin, Zhang, Shujian, Zhou, Wenxuan, Ghassemi, Marzyeh, Zhao, Sanqiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911604821458944
author Xiao, Yuxin
Zhang, Shujian
Zhou, Wenxuan
Ghassemi, Marzyeh
Zhao, Sanqiang
author_facet Xiao, Yuxin
Zhang, Shujian
Zhou, Wenxuan
Ghassemi, Marzyeh
Zhao, Sanqiang
contents To acquire instruction-following capabilities, large language models (LLMs) undergo instruction tuning, where they are trained on instruction-response pairs using next-token prediction (NTP). Efforts to improve instruction tuning often focus on higher-quality supervised fine-tuning (SFT) datasets, typically requiring data filtering with proprietary LLMs or human annotation. In this paper, we take a different approach by proposing SFTMix, a novel Mixup-based recipe that elevates LLM instruction tuning without relying on well-curated datasets. We observe that LLMs exhibit uneven confidence across the semantic representation space. We argue that examples with different confidence levels should play distinct roles in instruction tuning: Confident data is prone to overfitting, while unconfident data is harder to generalize. Based on this insight, SFTMix leverages training dynamics to identify examples with varying confidence levels. We then interpolate them to bridge the confidence gap and apply a Mixup-based regularization to support learning on these additional, interpolated examples. We demonstrate the effectiveness of SFTMix in both instruction-following and healthcare-specific SFT tasks, with consistent improvements across LLM families and SFT datasets of varying sizes and qualities. Extensive analyses across six directions highlight SFTMix's compatibility with data selection, adaptability to compute-constrained scenarios, and scalability to broader applications.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05248
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
Xiao, Yuxin
Zhang, Shujian
Zhou, Wenxuan
Ghassemi, Marzyeh
Zhao, Sanqiang
Computation and Language
Artificial Intelligence
Machine Learning
To acquire instruction-following capabilities, large language models (LLMs) undergo instruction tuning, where they are trained on instruction-response pairs using next-token prediction (NTP). Efforts to improve instruction tuning often focus on higher-quality supervised fine-tuning (SFT) datasets, typically requiring data filtering with proprietary LLMs or human annotation. In this paper, we take a different approach by proposing SFTMix, a novel Mixup-based recipe that elevates LLM instruction tuning without relying on well-curated datasets. We observe that LLMs exhibit uneven confidence across the semantic representation space. We argue that examples with different confidence levels should play distinct roles in instruction tuning: Confident data is prone to overfitting, while unconfident data is harder to generalize. Based on this insight, SFTMix leverages training dynamics to identify examples with varying confidence levels. We then interpolate them to bridge the confidence gap and apply a Mixup-based regularization to support learning on these additional, interpolated examples. We demonstrate the effectiveness of SFTMix in both instruction-following and healthcare-specific SFT tasks, with consistent improvements across LLM families and SFT datasets of varying sizes and qualities. Extensive analyses across six directions highlight SFTMix's compatibility with data selection, adaptability to compute-constrained scenarios, and scalability to broader applications.
title SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.05248