Towards Efficient Active Learning in NLP via Pretrained Representations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vysogorets, Artem, Gopal, Achintya
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911783650852864
author Vysogorets, Artem
Gopal, Achintya
author_facet Vysogorets, Artem
Gopal, Achintya
contents Fine-tuning Large Language Models (LLMs) is now a common approach for text classification in a wide range of applications. When labeled documents are scarce, active learning helps save annotation efforts but requires retraining of massive models on each acquisition iteration. We drastically expedite this process by using pretrained representations of LLMs within the active learning loop and, once the desired amount of labeled data is acquired, fine-tuning that or even a different pretrained LLM on this labeled data to achieve the best performance. As verified on common text classification benchmarks with pretrained BERT and RoBERTa as the backbone, our strategy yields similar performance to fine-tuning all the way through the active learning loop but is orders of magnitude less computationally expensive. The data acquired with our procedure generalizes across pretrained networks, allowing flexibility in choosing the final model or updating it as newer versions get released.
format Preprint
id arxiv_https___arxiv_org_abs_2402_15613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Efficient Active Learning in NLP via Pretrained Representations
Vysogorets, Artem
Gopal, Achintya
Machine Learning
Computation and Language
Fine-tuning Large Language Models (LLMs) is now a common approach for text classification in a wide range of applications. When labeled documents are scarce, active learning helps save annotation efforts but requires retraining of massive models on each acquisition iteration. We drastically expedite this process by using pretrained representations of LLMs within the active learning loop and, once the desired amount of labeled data is acquired, fine-tuning that or even a different pretrained LLM on this labeled data to achieve the best performance. As verified on common text classification benchmarks with pretrained BERT and RoBERTa as the backbone, our strategy yields similar performance to fine-tuning all the way through the active learning loop but is orders of magnitude less computationally expensive. The data acquired with our procedure generalizes across pretrained networks, allowing flexibility in choosing the final model or updating it as newer versions get released.
title Towards Efficient Active Learning in NLP via Pretrained Representations
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2402.15613