Relating Word Embedding Gender Biases to Gender Gaps: A Cross-Cultural Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Friedman, Scott, Schmer-Galunder, Sonja, Chen, Anthony, Rye, Jeffrey
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915752775254016
author Friedman, Scott
Schmer-Galunder, Sonja
Chen, Anthony
Rye, Jeffrey
author_facet Friedman, Scott
Schmer-Galunder, Sonja
Chen, Anthony
Rye, Jeffrey
contents Modern models for common NLP tasks often employ machine learning techniques and train on journalistic, social media, or other culturally-derived text. These have recently been scrutinized for racial and gender biases, rooting from inherent bias in their training text. These biases are often sub-optimal and recent work poses methods to rectify them; however, these biases may shed light on actual racial or gender gaps in the culture(s) that produced the training text, thereby helping us understand cultural context through big data. This paper presents an approach for quantifying gender bias in word embeddings, and then using them to characterize statistical gender gaps in education, politics, economics, and health. We validate these metrics on 2018 Twitter data spanning 51 U.S. regions and 99 countries. We correlate state and country word embedding biases with 18 international and 5 U.S.-based statistical gender gaps, characterizing regularities and predictive strength.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17203
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Relating Word Embedding Gender Biases to Gender Gaps: A Cross-Cultural Analysis
Friedman, Scott
Schmer-Galunder, Sonja
Chen, Anthony
Rye, Jeffrey
Computation and Language
68T50
I.2.7
Modern models for common NLP tasks often employ machine learning techniques and train on journalistic, social media, or other culturally-derived text. These have recently been scrutinized for racial and gender biases, rooting from inherent bias in their training text. These biases are often sub-optimal and recent work poses methods to rectify them; however, these biases may shed light on actual racial or gender gaps in the culture(s) that produced the training text, thereby helping us understand cultural context through big data. This paper presents an approach for quantifying gender bias in word embeddings, and then using them to characterize statistical gender gaps in education, politics, economics, and health. We validate these metrics on 2018 Twitter data spanning 51 U.S. regions and 99 countries. We correlate state and country word embedding biases with 18 international and 5 U.S.-based statistical gender gaps, characterizing regularities and predictive strength.
title Relating Word Embedding Gender Biases to Gender Gaps: A Cross-Cultural Analysis
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2601.17203