LlamBERT: Large-scale low-cost data annotation in NLP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Csanády, Bálint, Muzsai, Lajos, Vedres, Péter, Nádasdy, Zoltán, Lukács, András
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929286688014336
author Csanády, Bálint
Muzsai, Lajos
Vedres, Péter
Nádasdy, Zoltán
Lukács, András
author_facet Csanády, Bálint
Muzsai, Lajos
Vedres, Péter
Nádasdy, Zoltán
Lukács, András
contents Large Language Models (LLMs), such as GPT-4 and Llama 2, show remarkable proficiency in a wide range of natural language processing (NLP) tasks. Despite their effectiveness, the high costs associated with their use pose a challenge. We present LlamBERT, a hybrid approach that leverages LLMs to annotate a small subset of large, unlabeled databases and uses the results for fine-tuning transformer encoders like BERT and RoBERTa. This strategy is evaluated on two diverse datasets: the IMDb review dataset and the UMLS Meta-Thesaurus. Our results indicate that the LlamBERT approach slightly compromises on accuracy while offering much greater cost-effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15938
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LlamBERT: Large-scale low-cost data annotation in NLP
Csanády, Bálint
Muzsai, Lajos
Vedres, Péter
Nádasdy, Zoltán
Lukács, András
Computation and Language
Artificial Intelligence
Machine Learning
I.2.7; F.1.1
Large Language Models (LLMs), such as GPT-4 and Llama 2, show remarkable proficiency in a wide range of natural language processing (NLP) tasks. Despite their effectiveness, the high costs associated with their use pose a challenge. We present LlamBERT, a hybrid approach that leverages LLMs to annotate a small subset of large, unlabeled databases and uses the results for fine-tuning transformer encoders like BERT and RoBERTa. This strategy is evaluated on two diverse datasets: the IMDb review dataset and the UMLS Meta-Thesaurus. Our results indicate that the LlamBERT approach slightly compromises on accuracy while offering much greater cost-effectiveness.
title LlamBERT: Large-scale low-cost data annotation in NLP
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.7; F.1.1
url https://arxiv.org/abs/2403.15938