CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Vipul, Venkit, Pranav Narayanan, Laurençon, Hugo, Wilson, Shomir, Passonneau, Rebecca J.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916350028414976
author Gupta, Vipul
Venkit, Pranav Narayanan
Laurençon, Hugo
Wilson, Shomir
Passonneau, Rebecca J.
author_facet Gupta, Vipul
Venkit, Pranav Narayanan
Laurençon, Hugo
Wilson, Shomir
Passonneau, Rebecca J.
contents As language models (LMs) become increasingly powerful and widely used, it is important to quantify them for sociodemographic bias with potential for harm. Prior measures of bias are sensitive to perturbations in the templates designed to compare performance across social groups, due to factors such as low diversity or limited number of templates. Also, most previous work considers only one NLP task. We introduce Comprehensive Assessment of Language Models (CALM) for robust measurement of two types of universally relevant sociodemographic bias, gender and race. CALM integrates sixteen datasets for question-answering, sentiment analysis and natural language inference. Examples from each dataset are filtered to produce 224 templates with high diversity (e.g., length, vocabulary). We assemble 50 highly frequent person names for each of seven distinct demographic groups to generate 78,400 prompts covering the three NLP tasks. Our empirical evaluation shows that CALM bias scores are more robust and far less sensitive than previous bias measurements to perturbations in the templates, such as synonym substitution, or to random subset selection of templates. We apply CALM to 20 large language models, and find that for 2 language model series, larger parameter models tend to be more biased than smaller ones. The T0 series is the least biased model families, of the 20 LLMs investigated here. The code is available at https://github.com/vipulgupta1011/CALM.
format Preprint
id arxiv_https___arxiv_org_abs_2308_12539
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
Gupta, Vipul
Venkit, Pranav Narayanan
Laurençon, Hugo
Wilson, Shomir
Passonneau, Rebecca J.
Computation and Language
Artificial Intelligence
Machine Learning
As language models (LMs) become increasingly powerful and widely used, it is important to quantify them for sociodemographic bias with potential for harm. Prior measures of bias are sensitive to perturbations in the templates designed to compare performance across social groups, due to factors such as low diversity or limited number of templates. Also, most previous work considers only one NLP task. We introduce Comprehensive Assessment of Language Models (CALM) for robust measurement of two types of universally relevant sociodemographic bias, gender and race. CALM integrates sixteen datasets for question-answering, sentiment analysis and natural language inference. Examples from each dataset are filtered to produce 224 templates with high diversity (e.g., length, vocabulary). We assemble 50 highly frequent person names for each of seven distinct demographic groups to generate 78,400 prompts covering the three NLP tasks. Our empirical evaluation shows that CALM bias scores are more robust and far less sensitive than previous bias measurements to perturbations in the templates, such as synonym substitution, or to random subset selection of templates. We apply CALM to 20 large language models, and find that for 2 language model series, larger parameter models tend to be more biased than smaller ones. The T0 series is the least biased model families, of the 20 LLMs investigated here. The code is available at https://github.com/vipulgupta1011/CALM.
title CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2308.12539