BAID: A Benchmark for Bias Assessment of AI Detectors

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Basu, Priyam, Zhang, Yunfeng, Raheja, Vipul
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918246055149568
author Basu, Priyam
Zhang, Yunfeng
Raheja, Vipul
author_facet Basu, Priyam
Zhang, Yunfeng
Raheja, Vipul
contents AI-generated text detectors have recently gained adoption in educational and professional contexts. Prior research has uncovered isolated cases of bias, particularly against English Language Learners (ELLs) however, there is a lack of systematic evaluation of such systems across broader sociolinguistic factors. In this work, we propose BAID, a comprehensive evaluation framework for AI detectors across various types of biases. As a part of the framework, we introduce over 200k samples spanning 7 major categories: demographics, age, educational grade level, dialect, formality, political leaning, and topic. We also generated synthetic versions of each sample with carefully crafted prompts to preserve the original content while reflecting subgroup-specific writing styles. Using this, we evaluate four open-source state-of-the-art AI text detectors and find consistent disparities in detection performance, particularly low recall rates for texts from underrepresented groups. Our contributions provide a scalable, transparent approach for auditing AI detectors and emphasize the need for bias-aware evaluation before these tools are deployed for public use.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11505
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BAID: A Benchmark for Bias Assessment of AI Detectors
Basu, Priyam
Zhang, Yunfeng
Raheja, Vipul
Artificial Intelligence
Machine Learning
AI-generated text detectors have recently gained adoption in educational and professional contexts. Prior research has uncovered isolated cases of bias, particularly against English Language Learners (ELLs) however, there is a lack of systematic evaluation of such systems across broader sociolinguistic factors. In this work, we propose BAID, a comprehensive evaluation framework for AI detectors across various types of biases. As a part of the framework, we introduce over 200k samples spanning 7 major categories: demographics, age, educational grade level, dialect, formality, political leaning, and topic. We also generated synthetic versions of each sample with carefully crafted prompts to preserve the original content while reflecting subgroup-specific writing styles. Using this, we evaluate four open-source state-of-the-art AI text detectors and find consistent disparities in detection performance, particularly low recall rates for texts from underrepresented groups. Our contributions provide a scalable, transparent approach for auditing AI detectors and emphasize the need for bias-aware evaluation before these tools are deployed for public use.
title BAID: A Benchmark for Bias Assessment of AI Detectors
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.11505