Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Alipour, Shayan, Sen, Indira, Samory, Mattia, Mitra, Tanushree
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929600132546560
author Alipour, Shayan
Sen, Indira
Samory, Mattia
Mitra, Tanushree
author_facet Alipour, Shayan
Sen, Indira
Samory, Mattia
Mitra, Tanushree
contents Large language models (LLMs) are known to exhibit demographic biases, yet few studies systematically evaluate these biases across multiple datasets or account for confounding factors. In this work, we examine LLM alignment with human annotations in five offensive language datasets, comprising approximately 220K annotations. Our findings reveal that while demographic traits, particularly race, influence alignment, these effects are inconsistent across datasets and often entangled with other factors. Confounders -- such as document difficulty, annotator sensitivity, and within-group agreement -- account for more variation in alignment patterns than demographic traits alone. Specifically, alignment increases with higher annotator sensitivity and group agreement, while greater document difficulty corresponds to reduced alignment. Our results underscore the importance of multi-dataset analyses and confounder-aware methodologies in developing robust measures of demographic bias in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08977
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
Alipour, Shayan
Sen, Indira
Samory, Mattia
Mitra, Tanushree
Computers and Society
Computation and Language
Large language models (LLMs) are known to exhibit demographic biases, yet few studies systematically evaluate these biases across multiple datasets or account for confounding factors. In this work, we examine LLM alignment with human annotations in five offensive language datasets, comprising approximately 220K annotations. Our findings reveal that while demographic traits, particularly race, influence alignment, these effects are inconsistent across datasets and often entangled with other factors. Confounders -- such as document difficulty, annotator sensitivity, and within-group agreement -- account for more variation in alignment patterns than demographic traits alone. Specifically, alignment increases with higher annotator sensitivity and group agreement, while greater document difficulty corresponds to reduced alignment. Our results underscore the importance of multi-dataset analyses and confounder-aware methodologies in developing robust measures of demographic bias in LLMs.
title Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
topic Computers and Society
Computation and Language
url https://arxiv.org/abs/2411.08977