ProText: A benchmark dataset for measuring (mis)gendering in long-form texts
Fuente:
arXiv
Saved in:
| Main Authors: | Kotek, Hadas, Bowler, Margit, Sonnenberg, Patrick, Yang, Yu'an |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Protected group bias and stereotypes in Large Language Models
by: Kotek, Hadas, et al.
Published: (2024)
by: Kotek, Hadas, et al.
Published: (2024)
A large-scale image-text dataset benchmark for farmland segmentation
by: Tao, Chao, et al.
Published: (2025)
by: Tao, Chao, et al.
Published: (2025)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
by: Orgad, Hadas, et al.
Published: (2024)
by: Orgad, Hadas, et al.
Published: (2024)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
Frankentext: Stitching random text fragments into long-form narratives
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
AlleNoise: large-scale text classification benchmark dataset with real-world label noise
by: Rączkowska, Alicja, et al.
Published: (2024)
by: Rączkowska, Alicja, et al.
Published: (2024)
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
by: Song, Yixiao, et al.
Published: (2024)
by: Song, Yixiao, et al.
Published: (2024)
Stereotypical gender actions can be extracted from Web text
by: Herdağdelen, Amaç, et al.
Published: (2025)
by: Herdağdelen, Amaç, et al.
Published: (2025)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
by: Li, Kunning, et al.
Published: (2025)
by: Li, Kunning, et al.
Published: (2025)
CrowdCounter: A benchmark type-specific multi-target counterspeech dataset
by: Saha, Punyajoy, et al.
Published: (2024)
by: Saha, Punyajoy, et al.
Published: (2024)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
by: Arad, Dana, et al.
Published: (2023)
by: Arad, Dana, et al.
Published: (2023)
A multi-level multi-label text classification dataset of 19th century Ottoman and Russian literary and critical texts
by: Gokceoglu, Gokcen, et al.
Published: (2024)
by: Gokceoglu, Gokcen, et al.
Published: (2024)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
by: Cunha, Washington, et al.
Published: (2025)
by: Cunha, Washington, et al.
Published: (2025)
Detection of Adverse Drug Events in Dutch clinical free text documents using Transformer Models: benchmark study
by: Murphy, Rachel M., et al.
Published: (2025)
by: Murphy, Rachel M., et al.
Published: (2025)
CIDER: Context sensitive sentiment analysis for short-form text
by: Young, James C., et al.
Published: (2023)
by: Young, James C., et al.
Published: (2023)
Comparing representations of long clinical texts for the task of patient note-identification
by: Alsaidi, Safa, et al.
Published: (2025)
by: Alsaidi, Safa, et al.
Published: (2025)
Domain-specific long text classification from sparse relevant information
by: D'Cruz, Célia, et al.
Published: (2024)
by: D'Cruz, Célia, et al.
Published: (2024)
BeanCounter: A low-toxicity, large-scale, and open dataset of business-oriented text
by: Wang, Siyan, et al.
Published: (2024)
by: Wang, Siyan, et al.
Published: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
A dataset and benchmark for hospital course summarization with adapted large language models
by: Aali, Asad, et al.
Published: (2024)
by: Aali, Asad, et al.
Published: (2024)
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
by: Khallaf, Nouran, et al.
Published: (2025)
by: Khallaf, Nouran, et al.
Published: (2025)
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
IT5: Text-to-text Pretraining for Italian Language Understanding and Generation
by: Sarti, Gabriele, et al.
Published: (2022)
by: Sarti, Gabriele, et al.
Published: (2022)
ValiText -- a unified validation framework for computational text-based measures of social constructs
by: Birkenmaier, Lukas, et al.
Published: (2023)
by: Birkenmaier, Lukas, et al.
Published: (2023)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
Less than one percent of words would be affected by gender-inclusive language in German press texts
by: Müller-Spitzer, Carolin, et al.
Published: (2024)
by: Müller-Spitzer, Carolin, et al.
Published: (2024)
How do we measure privacy in text? A survey of text anonymization metrics
by: Ren, Yaxuan, et al.
Published: (2025)
by: Ren, Yaxuan, et al.
Published: (2025)
The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process
by: Contestabile, Matilde, et al.
Published: (2025)
by: Contestabile, Matilde, et al.
Published: (2025)
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries
by: Dong, Mingwen, et al.
Published: (2024)
by: Dong, Mingwen, et al.
Published: (2024)
TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models
by: Feng, Jiarui, et al.
Published: (2024)
by: Feng, Jiarui, et al.
Published: (2024)
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
by: Ging, Simon, et al.
Published: (2024)
by: Ging, Simon, et al.
Published: (2024)
A social context-aware graph-based multimodal attentive learning framework for disaster content classification during emergencies: a benchmark dataset and method
by: Dar, Shahid Shafi, et al.
Published: (2024)
by: Dar, Shahid Shafi, et al.
Published: (2024)
Text Difficulty Study: Do machines behave the same as humans regarding text difficulty?
by: Chen, Bowen, et al.
Published: (2022)
by: Chen, Bowen, et al.
Published: (2022)
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark
by: Wu, Jian, et al.
Published: (2024)
by: Wu, Jian, et al.
Published: (2024)
MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities
by: Durmaz, Ali Riza, et al.
Published: (2024)
by: Durmaz, Ali Riza, et al.
Published: (2024)
Synthetically generated text for supervised text analysis
by: Halterman, Andrew
Published: (2023)
by: Halterman, Andrew
Published: (2023)
Dynamic benchmarking framework for LLM-based conversational data capture
by: Aluffi, Pietro Alessandro, et al.
Published: (2025)
by: Aluffi, Pietro Alessandro, et al.
Published: (2025)
VeriFastScore: Speeding up long-form factuality evaluation
by: Rajendhran, Rishanth, et al.
Published: (2025)
by: Rajendhran, Rishanth, et al.
Published: (2025)
Designing large language model prompts to extract scores from messy text: A shared dataset and challenge
by: Thelwall, Mike
Published: (2026)
by: Thelwall, Mike
Published: (2026)
Similar Items
-
Protected group bias and stereotypes in Large Language Models
by: Kotek, Hadas, et al.
Published: (2024) -
A large-scale image-text dataset benchmark for farmland segmentation
by: Tao, Chao, et al.
Published: (2025) -
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
by: Orgad, Hadas, et al.
Published: (2024) -
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
by: He, Jie, et al.
Published: (2025) -
Frankentext: Stitching random text fragments into long-form narratives
by: Pham, Chau Minh, et al.
Published: (2025)