DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shin, Haebin, Ji, Lei, Liu, Xiao, Yu, Zhiwei, Chen, Qi, Gong, Yeyun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916902986579968
author Shin, Haebin
Ji, Lei
Liu, Xiao
Yu, Zhiwei
Chen, Qi
Gong, Yeyun
author_facet Shin, Haebin
Ji, Lei
Liu, Xiao
Yu, Zhiwei
Chen, Qi
Gong, Yeyun
contents As numerous instruction-tuning datasets continue to emerge during the post-training stage, dynamically balancing and optimizing their mixtures has become a critical challenge. To address this, we propose DynamixSFT, a dynamic and automated method for instruction-tuning dataset mixture optimization. We formulate the problem as a multi-armed bandit setup and introduce a Prior-scaled Boltzmann Exploration that softly anchors the updated sampling distribution to the original dataset proportions, thereby preserving the inherent diversity and coverage of the collection. Sampling probabilities are updated using a lightweight 1-Step Look-ahead Reward, reflecting how much the dataset contributes to improving the model's performance at its current state. When applied to the Tulu-v2-mixture collection comprising 16 instruction-tuning datasets, DynamixSFT achieves up to a 2.2% performance improvement across 10 benchmarks. Furthermore, we provide a comprehensive analysis and visualizations to offer deeper insights into the adaptive dynamics of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12116
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
Shin, Haebin
Ji, Lei
Liu, Xiao
Yu, Zhiwei
Chen, Qi
Gong, Yeyun
Machine Learning
Artificial Intelligence
Computation and Language
As numerous instruction-tuning datasets continue to emerge during the post-training stage, dynamically balancing and optimizing their mixtures has become a critical challenge. To address this, we propose DynamixSFT, a dynamic and automated method for instruction-tuning dataset mixture optimization. We formulate the problem as a multi-armed bandit setup and introduce a Prior-scaled Boltzmann Exploration that softly anchors the updated sampling distribution to the original dataset proportions, thereby preserving the inherent diversity and coverage of the collection. Sampling probabilities are updated using a lightweight 1-Step Look-ahead Reward, reflecting how much the dataset contributes to improving the model's performance at its current state. When applied to the Tulu-v2-mixture collection comprising 16 instruction-tuning datasets, DynamixSFT achieves up to a 2.2% performance improvement across 10 benchmarks. Furthermore, we provide a comprehensive analysis and visualizations to offer deeper insights into the adaptive dynamics of our method.
title DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.12116