Compute-Constrained Data Selection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yin, Junjie Oscar, Rush, Alexander M.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910905547096064
author Yin, Junjie Oscar
Rush, Alexander M.
author_facet Yin, Junjie Oscar
Rush, Alexander M.
contents Data selection can reduce the amount of training data needed to finetune LLMs; however, the efficacy of data selection scales directly with its compute. Motivated by the practical challenge of compute-constrained finetuning, we consider the setting in which both the cost of selecting data and training are budgeted for. We first formalize the problem of data selection with a cost-aware utility function, and model the data selection problem as trading off initial-selection cost for training gain. We run a comprehensive sweep of experiments across multiple tasks, varying compute budget by scaling finetuning tokens, model sizes, and data selection compute. Interestingly we find that many powerful data selection methods are almost never compute-optimal, and that cheaper data selection alternatives dominate both from a theoretical and empirical perspective. For compute-optimal training, we find that perplexity and gradient data selection require training-to-selection model size ratios of 5x and 10x, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16208
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Compute-Constrained Data Selection
Yin, Junjie Oscar
Rush, Alexander M.
Machine Learning
Artificial Intelligence
Computation and Language
Data selection can reduce the amount of training data needed to finetune LLMs; however, the efficacy of data selection scales directly with its compute. Motivated by the practical challenge of compute-constrained finetuning, we consider the setting in which both the cost of selecting data and training are budgeted for. We first formalize the problem of data selection with a cost-aware utility function, and model the data selection problem as trading off initial-selection cost for training gain. We run a comprehensive sweep of experiments across multiple tasks, varying compute budget by scaling finetuning tokens, model sizes, and data selection compute. Interestingly we find that many powerful data selection methods are almost never compute-optimal, and that cheaper data selection alternatives dominate both from a theoretical and empirical perspective. For compute-optimal training, we find that perplexity and gradient data selection require training-to-selection model size ratios of 5x and 10x, respectively.
title Compute-Constrained Data Selection
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.16208