GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Felkner, Virginia K., Thompson, Jennifer A., May, Jonathan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914810423148544
author Felkner, Virginia K.
Thompson, Jennifer A.
May, Jonathan
author_facet Felkner, Virginia K.
Thompson, Jennifer A.
May, Jonathan
contents Social biases in LLMs are usually measured via bias benchmark datasets. Current benchmarks have limitations in scope, grounding, quality, and human effort required. Previous work has shown success with a community-sourced, rather than crowd-sourced, approach to benchmark development. However, this work still required considerable effort from annotators with relevant lived experience. This paper explores whether an LLM (specifically, GPT-3.5-Turbo) can assist with the task of developing a bias benchmark dataset from responses to an open-ended community survey. We also extend the previous work to a new community and set of biases: the Jewish community and antisemitism. Our analysis shows that GPT-3.5-Turbo has poor performance on this annotation task and produces unacceptable quality issues in its output. Thus, we conclude that GPT-3.5-Turbo is not an appropriate substitute for human annotation in sensitive tasks related to social biases, and that its use actually negates many of the benefits of community-sourcing bias benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15760
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
Felkner, Virginia K.
Thompson, Jennifer A.
May, Jonathan
Computation and Language
Computers and Society
I.2.7; K.4.2
Social biases in LLMs are usually measured via bias benchmark datasets. Current benchmarks have limitations in scope, grounding, quality, and human effort required. Previous work has shown success with a community-sourced, rather than crowd-sourced, approach to benchmark development. However, this work still required considerable effort from annotators with relevant lived experience. This paper explores whether an LLM (specifically, GPT-3.5-Turbo) can assist with the task of developing a bias benchmark dataset from responses to an open-ended community survey. We also extend the previous work to a new community and set of biases: the Jewish community and antisemitism. Our analysis shows that GPT-3.5-Turbo has poor performance on this annotation task and produces unacceptable quality issues in its output. Thus, we conclude that GPT-3.5-Turbo is not an appropriate substitute for human annotation in sensitive tasks related to social biases, and that its use actually negates many of the benefits of community-sourcing bias benchmarks.
title GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
topic Computation and Language
Computers and Society
I.2.7; K.4.2
url https://arxiv.org/abs/2405.15760