Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hübotter, Jonas, Bongni, Sascha, Hakimi, Ido, Krause, Andreas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917917176627200
author Hübotter, Jonas
Bongni, Sascha
Hakimi, Ido
Krause, Andreas
author_facet Hübotter, Jonas
Bongni, Sascha
Hakimi, Ido
Krause, Andreas
contents Recent efforts in fine-tuning language models often rely on automatic data selection, commonly using Nearest Neighbors retrieval from large datasets. However, we theoretically show that this approach tends to select redundant data, limiting its effectiveness or even hurting performance. To address this, we introduce SIFT, a data selection algorithm designed to reduce uncertainty about the model's response given a prompt, which unifies ideas from retrieval and active learning. Whereas Nearest Neighbor retrieval typically fails in the presence of information duplication, SIFT accounts for information duplication and optimizes the overall information gain of the selected examples. We focus our evaluations on fine-tuning at test-time for prompt-specific language modeling on the Pile dataset, and show that SIFT consistently outperforms Nearest Neighbor retrieval, with minimal computational overhead. Moreover, we show that our uncertainty estimates can predict the performance gain of test-time fine-tuning, and use this to develop an adaptive algorithm that invests test-time compute proportional to realized performance gains. We provide the $\texttt{activeft}$ (Active Fine-Tuning) library which can be used as a drop-in replacement for Nearest Neighbor retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08020
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
Hübotter, Jonas
Bongni, Sascha
Hakimi, Ido
Krause, Andreas
Machine Learning
Artificial Intelligence
Recent efforts in fine-tuning language models often rely on automatic data selection, commonly using Nearest Neighbors retrieval from large datasets. However, we theoretically show that this approach tends to select redundant data, limiting its effectiveness or even hurting performance. To address this, we introduce SIFT, a data selection algorithm designed to reduce uncertainty about the model's response given a prompt, which unifies ideas from retrieval and active learning. Whereas Nearest Neighbor retrieval typically fails in the presence of information duplication, SIFT accounts for information duplication and optimizes the overall information gain of the selected examples. We focus our evaluations on fine-tuning at test-time for prompt-specific language modeling on the Pile dataset, and show that SIFT consistently outperforms Nearest Neighbor retrieval, with minimal computational overhead. Moreover, we show that our uncertainty estimates can predict the performance gain of test-time fine-tuning, and use this to develop an adaptive algorithm that invests test-time compute proportional to realized performance gains. We provide the $\texttt{activeft}$ (Active Fine-Tuning) library which can be used as a drop-in replacement for Nearest Neighbor retrieval.
title Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.08020