Identifying Fairness Issues in Automatically Generated Testing Content

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stowe, Kevin, Longwill, Benny, Francis, Alyssa, Aoyama, Tatsuya, Ghosh, Debanjan, Somasundaran, Swapna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911862476505088
author Stowe, Kevin
Longwill, Benny
Francis, Alyssa
Aoyama, Tatsuya
Ghosh, Debanjan
Somasundaran, Swapna
author_facet Stowe, Kevin
Longwill, Benny
Francis, Alyssa
Aoyama, Tatsuya
Ghosh, Debanjan
Somasundaran, Swapna
contents Natural language generation tools are powerful and effective for generating content. However, language models are known to display bias and fairness issues, making them impractical to deploy for many use cases. We here focus on how fairness issues impact automatically generated test content, which can have stringent requirements to ensure the test measures only what it was intended to measure. Specifically, we review test content generated for a large-scale standardized English proficiency test with the goal of identifying content that only pertains to a certain subset of the test population as well as content that has the potential to be upsetting or distracting to some test takers. Issues like these could inadvertently impact a test taker's score and thus should be avoided. This kind of content does not reflect the more commonly-acknowledged biases, making it challenging even for modern models that contain safeguards. We build a dataset of 601 generated texts annotated for fairness and explore a variety of methods for classification: fine-tuning, topic-based classification, and prompting, including few-shot and self-correcting prompts. We find that combining prompt self-correction and few-shot learning performs best, yielding an F1 score of 0.79 on our held-out test set, while much smaller BERT- and topic-based models have competitive performance on out-of-domain data.
format Preprint
id arxiv_https___arxiv_org_abs_2404_15104
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Identifying Fairness Issues in Automatically Generated Testing Content
Stowe, Kevin
Longwill, Benny
Francis, Alyssa
Aoyama, Tatsuya
Ghosh, Debanjan
Somasundaran, Swapna
Computation and Language
I.2.7
Natural language generation tools are powerful and effective for generating content. However, language models are known to display bias and fairness issues, making them impractical to deploy for many use cases. We here focus on how fairness issues impact automatically generated test content, which can have stringent requirements to ensure the test measures only what it was intended to measure. Specifically, we review test content generated for a large-scale standardized English proficiency test with the goal of identifying content that only pertains to a certain subset of the test population as well as content that has the potential to be upsetting or distracting to some test takers. Issues like these could inadvertently impact a test taker's score and thus should be avoided. This kind of content does not reflect the more commonly-acknowledged biases, making it challenging even for modern models that contain safeguards. We build a dataset of 601 generated texts annotated for fairness and explore a variety of methods for classification: fine-tuning, topic-based classification, and prompting, including few-shot and self-correcting prompts. We find that combining prompt self-correction and few-shot learning performs best, yielding an F1 score of 0.79 on our held-out test set, while much smaller BERT- and topic-based models have competitive performance on out-of-domain data.
title Identifying Fairness Issues in Automatically Generated Testing Content
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2404.15104