LlamBERT: Large-scale low-cost data annotation in NLP
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929286688014336 |
|---|---|
| author | Csanády, Bálint Muzsai, Lajos Vedres, Péter Nádasdy, Zoltán Lukács, András |
| author_facet | Csanády, Bálint Muzsai, Lajos Vedres, Péter Nádasdy, Zoltán Lukács, András |
| contents | Large Language Models (LLMs), such as GPT-4 and Llama 2, show remarkable proficiency in a wide range of natural language processing (NLP) tasks. Despite their effectiveness, the high costs associated with their use pose a challenge. We present LlamBERT, a hybrid approach that leverages LLMs to annotate a small subset of large, unlabeled databases and uses the results for fine-tuning transformer encoders like BERT and RoBERTa. This strategy is evaluated on two diverse datasets: the IMDb review dataset and the UMLS Meta-Thesaurus. Our results indicate that the LlamBERT approach slightly compromises on accuracy while offering much greater cost-effectiveness. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_15938 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LlamBERT: Large-scale low-cost data annotation in NLP Csanády, Bálint Muzsai, Lajos Vedres, Péter Nádasdy, Zoltán Lukács, András Computation and Language Artificial Intelligence Machine Learning I.2.7; F.1.1 Large Language Models (LLMs), such as GPT-4 and Llama 2, show remarkable proficiency in a wide range of natural language processing (NLP) tasks. Despite their effectiveness, the high costs associated with their use pose a challenge. We present LlamBERT, a hybrid approach that leverages LLMs to annotate a small subset of large, unlabeled databases and uses the results for fine-tuning transformer encoders like BERT and RoBERTa. This strategy is evaluated on two diverse datasets: the IMDb review dataset and the UMLS Meta-Thesaurus. Our results indicate that the LlamBERT approach slightly compromises on accuracy while offering much greater cost-effectiveness. |
| title | LlamBERT: Large-scale low-cost data annotation in NLP |
| topic | Computation and Language Artificial Intelligence Machine Learning I.2.7; F.1.1 |
| url | https://arxiv.org/abs/2403.15938 |