Saved in:
Bibliographic Details
Main Authors: Barrie, Christopher, Palaiologou, Elli, Törnberg, Petter
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.02039
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914568099332096
author Barrie, Christopher
Palaiologou, Elli
Törnberg, Petter
author_facet Barrie, Christopher
Palaiologou, Elli
Törnberg, Petter
contents Researchers are increasingly using language models (LMs) for text annotation. These approaches rely only on a prompt telling the model to return a given output according to a set of instructions. The reproducibility of LM outputs may nonetheless be vulnerable to small changes in the prompt design. This calls into question the replicability of classification routines. To tackle this problem, researchers have typically tested a variety of semantically similar prompts to determine what we call ``prompt stability." These approaches remain ad-hoc and task specific. In this article, we propose a general framework for diagnosing prompt stability by adapting traditional approaches to intra- and inter-coder reliability scoring. We call the resulting metric the Prompt Stability Score (PSS) and provide a Python package \texttt{promptstability} for its estimation. Using six different datasets and twelve outcomes, we classify $\sim$3.1m rows of data and $\sim$300m input tokens to: a) diagnose when prompt stability is low; and b) demonstrate the functionality of the package. We conclude by providing best practice recommendations for applied researchers.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prompt Stability Scoring for Text Annotation with Large Language Models
Barrie, Christopher
Palaiologou, Elli
Törnberg, Petter
Computation and Language
Researchers are increasingly using language models (LMs) for text annotation. These approaches rely only on a prompt telling the model to return a given output according to a set of instructions. The reproducibility of LM outputs may nonetheless be vulnerable to small changes in the prompt design. This calls into question the replicability of classification routines. To tackle this problem, researchers have typically tested a variety of semantically similar prompts to determine what we call ``prompt stability." These approaches remain ad-hoc and task specific. In this article, we propose a general framework for diagnosing prompt stability by adapting traditional approaches to intra- and inter-coder reliability scoring. We call the resulting metric the Prompt Stability Score (PSS) and provide a Python package \texttt{promptstability} for its estimation. Using six different datasets and twelve outcomes, we classify $\sim$3.1m rows of data and $\sim$300m input tokens to: a) diagnose when prompt stability is low; and b) demonstrate the functionality of the package. We conclude by providing best practice recommendations for applied researchers.
title Prompt Stability Scoring for Text Annotation with Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2407.02039