An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhatt, Gantavya, Chen, Yifang, Das, Arnav M., Zhang, Jifan, Truong, Sang T., Mussmann, Stephen, Zhu, Yinglun, Bilmes, Jeffrey, Du, Simon S., Jamieson, Kevin, Ash, Jordan T., Nowak, Robert D.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910516969996288
author Bhatt, Gantavya
Chen, Yifang
Das, Arnav M.
Zhang, Jifan
Truong, Sang T.
Mussmann, Stephen
Zhu, Yinglun
Bilmes, Jeffrey
Du, Simon S.
Jamieson, Kevin
Ash, Jordan T.
Nowak, Robert D.
author_facet Bhatt, Gantavya
Chen, Yifang
Das, Arnav M.
Zhang, Jifan
Truong, Sang T.
Mussmann, Stephen
Zhu, Yinglun
Bilmes, Jeffrey
Du, Simon S.
Jamieson, Kevin
Ash, Jordan T.
Nowak, Robert D.
contents Supervised finetuning (SFT) on instruction datasets has played a crucial role in achieving the remarkable zero-shot generalization capabilities observed in modern large language models (LLMs). However, the annotation efforts required to produce high quality responses for instructions are becoming prohibitively expensive, especially as the number of tasks spanned by instruction datasets continues to increase. Active learning is effective in identifying useful subsets of samples to annotate from an unlabeled pool, but its high computational cost remains a barrier to its widespread applicability in the context of LLMs. To mitigate the annotation cost of SFT and circumvent the computational bottlenecks of active learning, we propose using experimental design. Experimental design techniques select the most informative samples to label, and typically maximize some notion of uncertainty and/or diversity. In our work, we implement a framework that evaluates several existing and novel experimental design techniques and find that these methods consistently yield significant gains in label efficiency with little computational overhead. On generative tasks, our methods achieve the same generalization performance with only $50\%$ of annotation cost required by random sampling.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06692
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models
Bhatt, Gantavya
Chen, Yifang
Das, Arnav M.
Zhang, Jifan
Truong, Sang T.
Mussmann, Stephen
Zhu, Yinglun
Bilmes, Jeffrey
Du, Simon S.
Jamieson, Kevin
Ash, Jordan T.
Nowak, Robert D.
Computation and Language
Artificial Intelligence
Machine Learning
Supervised finetuning (SFT) on instruction datasets has played a crucial role in achieving the remarkable zero-shot generalization capabilities observed in modern large language models (LLMs). However, the annotation efforts required to produce high quality responses for instructions are becoming prohibitively expensive, especially as the number of tasks spanned by instruction datasets continues to increase. Active learning is effective in identifying useful subsets of samples to annotate from an unlabeled pool, but its high computational cost remains a barrier to its widespread applicability in the context of LLMs. To mitigate the annotation cost of SFT and circumvent the computational bottlenecks of active learning, we propose using experimental design. Experimental design techniques select the most informative samples to label, and typically maximize some notion of uncertainty and/or diversity. In our work, we implement a framework that evaluates several existing and novel experimental design techniques and find that these methods consistently yield significant gains in label efficiency with little computational overhead. On generative tasks, our methods achieve the same generalization performance with only $50\%$ of annotation cost required by random sampling.
title An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2401.06692