ProText: A benchmark dataset for measuring (mis)gendering in long-form texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kotek, Hadas, Bowler, Margit, Sonnenberg, Patrick, Yang, Yu'an
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908920135548928
author Kotek, Hadas
Bowler, Margit
Sonnenberg, Patrick
Yang, Yu'an
author_facet Kotek, Hadas
Bowler, Margit
Sonnenberg, Patrick
Yang, Yu'an
contents We introduce ProText, a dataset for measuring gendering and misgendering in stylistically diverse long-form English texts. ProText spans three dimensions: Theme nouns (names, occupations, titles, kinship terms), Theme category (stereotypically male, stereotypically female, gender-neutral/non-gendered), and Pronoun category (masculine, feminine, gender-neutral, none). The dataset is designed to probe (mis)gendering in text transformations such as summarization and rewrites using state-of-the-art Large Language Models, extending beyond traditional pronoun resolution benchmarks and beyond the gender binary. We validated ProText through a mini case study, showing that even with just two prompts and two models, we can draw nuanced insights regarding gender bias, stereotyping, misgendering, and gendering. We reveal systematic gender bias, particularly when inputs contain no explicit gender cues or when models default to heteronormative assumptions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27838
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProText: A benchmark dataset for measuring (mis)gendering in long-form texts
Kotek, Hadas
Bowler, Margit
Sonnenberg, Patrick
Yang, Yu'an
Computation and Language
We introduce ProText, a dataset for measuring gendering and misgendering in stylistically diverse long-form English texts. ProText spans three dimensions: Theme nouns (names, occupations, titles, kinship terms), Theme category (stereotypically male, stereotypically female, gender-neutral/non-gendered), and Pronoun category (masculine, feminine, gender-neutral, none). The dataset is designed to probe (mis)gendering in text transformations such as summarization and rewrites using state-of-the-art Large Language Models, extending beyond traditional pronoun resolution benchmarks and beyond the gender binary. We validated ProText through a mini case study, showing that even with just two prompts and two models, we can draw nuanced insights regarding gender bias, stereotyping, misgendering, and gendering. We reveal systematic gender bias, particularly when inputs contain no explicit gender cues or when models default to heteronormative assumptions.
title ProText: A benchmark dataset for measuring (mis)gendering in long-form texts
topic Computation and Language
url https://arxiv.org/abs/2603.27838