Enhancing Low-Resource LLMs Classification with PEFT and Synthetic Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Patwa, Parth, Filice, Simone, Chen, Zhiyu, Castellucci, Giuseppe, Rokhlenko, Oleg, Malmasi, Shervin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913297100439552
author Patwa, Parth
Filice, Simone
Chen, Zhiyu
Castellucci, Giuseppe
Rokhlenko, Oleg
Malmasi, Shervin
author_facet Patwa, Parth
Filice, Simone
Chen, Zhiyu
Castellucci, Giuseppe
Rokhlenko, Oleg
Malmasi, Shervin
contents Large Language Models (LLMs) operating in 0-shot or few-shot settings achieve competitive results in Text Classification tasks. In-Context Learning (ICL) typically achieves better accuracy than the 0-shot setting, but it pays in terms of efficiency, due to the longer input prompt. In this paper, we propose a strategy to make LLMs as efficient as 0-shot text classifiers, while getting comparable or better accuracy than ICL. Our solution targets the low resource setting, i.e., when only 4 examples per class are available. Using a single LLM and few-shot real data we perform a sequence of generation, filtering and Parameter-Efficient Fine-Tuning steps to create a robust and efficient classifier. Experimental results show that our approach leads to competitive results on multiple text classification datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02422
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Low-Resource LLMs Classification with PEFT and Synthetic Data
Patwa, Parth
Filice, Simone
Chen, Zhiyu
Castellucci, Giuseppe
Rokhlenko, Oleg
Malmasi, Shervin
Computation and Language
Machine Learning
Large Language Models (LLMs) operating in 0-shot or few-shot settings achieve competitive results in Text Classification tasks. In-Context Learning (ICL) typically achieves better accuracy than the 0-shot setting, but it pays in terms of efficiency, due to the longer input prompt. In this paper, we propose a strategy to make LLMs as efficient as 0-shot text classifiers, while getting comparable or better accuracy than ICL. Our solution targets the low resource setting, i.e., when only 4 examples per class are available. Using a single LLM and few-shot real data we perform a sequence of generation, filtering and Parameter-Efficient Fine-Tuning steps to create a robust and efficient classifier. Experimental results show that our approach leads to competitive results on multiple text classification datasets.
title Enhancing Low-Resource LLMs Classification with PEFT and Synthetic Data
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2404.02422