Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vargas, Francisco, Cotterell, Ryan
Format: Preprint
Published: 2020
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914804855209984
author Vargas, Francisco
Cotterell, Ryan
author_facet Vargas, Francisco
Cotterell, Ryan
contents Bolukbasi et al. (2016) presents one of the first gender bias mitigation techniques for word representations. Their method takes pre-trained word representations as input and attempts to isolate a linear subspace that captures most of the gender bias in the representations. As judged by an analogical evaluation task, their method virtually eliminates gender bias in the representations. However, an implicit and untested assumption of their method is that the bias subspace is actually linear. In this work, we generalize their method to a kernelized, nonlinear version. We take inspiration from kernel principal component analysis and derive a nonlinear bias isolation technique. We discuss and overcome some of the practical drawbacks of our method for non-linear gender bias mitigation in word representations and analyze empirically whether the bias subspace is actually linear. Our analysis shows that gender bias is in fact well captured by a linear subspace, justifying the assumption of Bolukbasi et al. (2016).
format Preprint
id arxiv_https___arxiv_org_abs_2009_09435
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
Vargas, Francisco
Cotterell, Ryan
Machine Learning
Computation and Language
Computers and Society
Bolukbasi et al. (2016) presents one of the first gender bias mitigation techniques for word representations. Their method takes pre-trained word representations as input and attempts to isolate a linear subspace that captures most of the gender bias in the representations. As judged by an analogical evaluation task, their method virtually eliminates gender bias in the representations. However, an implicit and untested assumption of their method is that the bias subspace is actually linear. In this work, we generalize their method to a kernelized, nonlinear version. We take inspiration from kernel principal component analysis and derive a nonlinear bias isolation technique. We discuss and overcome some of the practical drawbacks of our method for non-linear gender bias mitigation in word representations and analyze empirically whether the bias subspace is actually linear. Our analysis shows that gender bias is in fact well captured by a linear subspace, justifying the assumption of Bolukbasi et al. (2016).
title Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
topic Machine Learning
Computation and Language
Computers and Society
url https://arxiv.org/abs/2009.09435