FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Athena, Patil, Tanush, Saxena, Ansh, Fu, Yicheng, O'Brien, Sean, Zhu, Kevin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917974110109696
author Wen, Athena
Patil, Tanush
Saxena, Ansh
Fu, Yicheng
O'Brien, Sean
Zhu, Kevin
author_facet Wen, Athena
Patil, Tanush
Saxena, Ansh
Fu, Yicheng
O'Brien, Sean
Zhu, Kevin
contents In an era where AI-driven hiring is transforming recruitment practices, concerns about fairness and bias have become increasingly important. To explore these issues, we introduce a benchmark, FAIRE (Fairness Assessment In Resume Evaluation), to test for racial and gender bias in large language models (LLMs) used to evaluate resumes across different industries. We use two methods-direct scoring and ranking-to measure how model performance changes when resumes are slightly altered to reflect different racial or gender identities. Our findings reveal that while every model exhibits some degree of bias, the magnitude and direction vary considerably. This benchmark provides a clear way to examine these differences and offers valuable insights into the fairness of AI-based hiring tools. It highlights the urgent need for strategies to reduce bias in AI-driven recruitment. Our benchmark code and dataset are open-sourced at our repository: https://github.com/athenawen/FAIRE-Fairness-Assessment-In-Resume-Evaluation.git.
format Preprint
id arxiv_https___arxiv_org_abs_2504_01420
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations
Wen, Athena
Patil, Tanush
Saxena, Ansh
Fu, Yicheng
O'Brien, Sean
Zhu, Kevin
Computation and Language
Artificial Intelligence
In an era where AI-driven hiring is transforming recruitment practices, concerns about fairness and bias have become increasingly important. To explore these issues, we introduce a benchmark, FAIRE (Fairness Assessment In Resume Evaluation), to test for racial and gender bias in large language models (LLMs) used to evaluate resumes across different industries. We use two methods-direct scoring and ranking-to measure how model performance changes when resumes are slightly altered to reflect different racial or gender identities. Our findings reveal that while every model exhibits some degree of bias, the magnitude and direction vary considerably. This benchmark provides a clear way to examine these differences and offers valuable insights into the fairness of AI-based hiring tools. It highlights the urgent need for strategies to reduce bias in AI-driven recruitment. Our benchmark code and dataset are open-sourced at our repository: https://github.com/athenawen/FAIRE-Fairness-Assessment-In-Resume-Evaluation.git.
title FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.01420