Political DEBATE: Efficient Zero-shot and Few-shot Classifiers for Political Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Burnham, Michael, Kahn, Kayla, Wang, Ryan Yank, Peng, Rachel X.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909989152489472
author Burnham, Michael
Kahn, Kayla
Wang, Ryan Yank
Peng, Rachel X.
author_facet Burnham, Michael
Kahn, Kayla
Wang, Ryan Yank
Peng, Rachel X.
contents Social scientists quickly adopted large language models due to their ability to annotate documents without supervised training, an ability known as zero-shot learning. However, due to their compute demands, cost, and often proprietary nature, these models are often at odds with replication and open science standards. This paper introduces the Political DEBATE (DeBERTa Algorithm for Textual Entailment) language models for zero-shot and few-shot classification of political documents. These models are not only as good, or better than, state-of-the art large language models at zero and few-shot classification, but are orders of magnitude more efficient and completely open source. By training the models on a simple random sample of 10-25 documents, they can outperform supervised classifiers trained on hundreds or thousands of documents and state-of-the-art generative models with complex, engineered prompts. Additionally, we release the PolNLI dataset used to train these models -- a corpus of over 200,000 political documents with highly accurate labels across over 800 classification tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2409_02078
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Political DEBATE: Efficient Zero-shot and Few-shot Classifiers for Political Text
Burnham, Michael
Kahn, Kayla
Wang, Ryan Yank
Peng, Rachel X.
Computation and Language
Social scientists quickly adopted large language models due to their ability to annotate documents without supervised training, an ability known as zero-shot learning. However, due to their compute demands, cost, and often proprietary nature, these models are often at odds with replication and open science standards. This paper introduces the Political DEBATE (DeBERTa Algorithm for Textual Entailment) language models for zero-shot and few-shot classification of political documents. These models are not only as good, or better than, state-of-the art large language models at zero and few-shot classification, but are orders of magnitude more efficient and completely open source. By training the models on a simple random sample of 10-25 documents, they can outperform supervised classifiers trained on hundreds or thousands of documents and state-of-the-art generative models with complex, engineered prompts. Additionally, we release the PolNLI dataset used to train these models -- a corpus of over 200,000 political documents with highly accurate labels across over 800 classification tasks.
title Political DEBATE: Efficient Zero-shot and Few-shot Classifiers for Political Text
topic Computation and Language
url https://arxiv.org/abs/2409.02078