Logistic Regression makes small LLMs strong and explainable "tens-of-shot" classifiers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Buckmann, Marcus, Hill, Edward
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917794101067776
author Buckmann, Marcus
Hill, Edward
author_facet Buckmann, Marcus
Hill, Edward
contents For simple classification tasks, we show that users can benefit from the advantages of using small, local, generative language models instead of large commercial models without a trade-off in performance or introducing extra labelling costs. These advantages, including those around privacy, availability, cost, and explainability, are important both in commercial applications and in the broader democratisation of AI. Through experiments on 17 sentence classification tasks (2-4 classes), we show that penalised logistic regression on the embeddings from a small LLM equals (and usually betters) the performance of a large LLM in the "tens-of-shot" regime. This requires no more labelled instances than are needed to validate the performance of the large LLM. Finally, we extract stable and sensible explanations for classification decisions.
format Preprint
id arxiv_https___arxiv_org_abs_2408_03414
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Logistic Regression makes small LLMs strong and explainable "tens-of-shot" classifiers
Buckmann, Marcus
Hill, Edward
Computation and Language
Machine Learning
68T50 (Primary), 62J07 (Secondary)
I.2.7
For simple classification tasks, we show that users can benefit from the advantages of using small, local, generative language models instead of large commercial models without a trade-off in performance or introducing extra labelling costs. These advantages, including those around privacy, availability, cost, and explainability, are important both in commercial applications and in the broader democratisation of AI. Through experiments on 17 sentence classification tasks (2-4 classes), we show that penalised logistic regression on the embeddings from a small LLM equals (and usually betters) the performance of a large LLM in the "tens-of-shot" regime. This requires no more labelled instances than are needed to validate the performance of the large LLM. Finally, we extract stable and sensible explanations for classification decisions.
title Logistic Regression makes small LLMs strong and explainable "tens-of-shot" classifiers
topic Computation and Language
Machine Learning
68T50 (Primary), 62J07 (Secondary)
I.2.7
url https://arxiv.org/abs/2408.03414