Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zou, Heming, Mao, Yixiu, Qu, Yun, Wang, Qi, Ji, Xiangyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917237743419392
author Zou, Heming
Mao, Yixiu
Qu, Yun
Wang, Qi
Ji, Xiangyang
author_facet Zou, Heming
Mao, Yixiu
Qu, Yun
Wang, Qi
Ji, Xiangyang
contents Supervised fine-tuning (SFT) is a commonly used technique to adapt large language models (LLMs) to downstream tasks. In practice, SFT on a full dataset is computationally expensive and sometimes suffers from overfitting or bias amplification. This facilitates the rise of data curation in SFT, which prioritizes the most valuable data to optimze. This work studies the online batch selection family that dynamically scores and filters samples during the training process. However, existing popular methods often (i) rely merely on the utility of data to select a subset while neglecting other crucial factors like diversity, (ii) rely on external resources such as reference models or validation sets, and (iii) incur extra training time over full-dataset training. To address these limitations, this work develops UDS (Utility-Diversity Sampling), a framework for efficient online batch selection in SFT. UDS leverages the nuclear norm of the logits matrix to capture both data utility and intra-sample diversity, while estimating inter-sample diversity through efficient low-dimensional embedding comparisons with a lightweight memory buffer of historical samples. Such a design eliminates the need for external resources and unnecessary backpropagation, securing computational efficiency. Experiments on multiple benchmarks demonstrate that UDS consistently outperforms state-of-the-art online batch selection methods under varying data budgets, and significantly reduces training time compared to full-dataset fine-tuning. Code is available at https://github.com/gfyddha/UDS.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16882
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
Zou, Heming
Mao, Yixiu
Qu, Yun
Wang, Qi
Ji, Xiangyang
Machine Learning
Artificial Intelligence
Computation and Language
Supervised fine-tuning (SFT) is a commonly used technique to adapt large language models (LLMs) to downstream tasks. In practice, SFT on a full dataset is computationally expensive and sometimes suffers from overfitting or bias amplification. This facilitates the rise of data curation in SFT, which prioritizes the most valuable data to optimze. This work studies the online batch selection family that dynamically scores and filters samples during the training process. However, existing popular methods often (i) rely merely on the utility of data to select a subset while neglecting other crucial factors like diversity, (ii) rely on external resources such as reference models or validation sets, and (iii) incur extra training time over full-dataset training. To address these limitations, this work develops UDS (Utility-Diversity Sampling), a framework for efficient online batch selection in SFT. UDS leverages the nuclear norm of the logits matrix to capture both data utility and intra-sample diversity, while estimating inter-sample diversity through efficient low-dimensional embedding comparisons with a lightweight memory buffer of historical samples. Such a design eliminates the need for external resources and unnecessary backpropagation, securing computational efficiency. Experiments on multiple benchmarks demonstrate that UDS consistently outperforms state-of-the-art online batch selection methods under varying data budgets, and significantly reduces training time compared to full-dataset fine-tuning. Code is available at https://github.com/gfyddha/UDS.
title Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.16882