Collapsed Language Models Promote Fairness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Jingxuan, Chen, Wuyang, Li, Linyi, Zhao, Yao, Wei, Yunchao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908204213993472
author Xu, Jingxuan
Chen, Wuyang
Li, Linyi
Zhao, Yao
Wei, Yunchao
author_facet Xu, Jingxuan
Chen, Wuyang
Li, Linyi
Zhao, Yao
Wei, Yunchao
contents To mitigate societal biases implicitly encoded in recent successful pretrained language models, a diverse array of approaches have been proposed to encourage model fairness, focusing on prompting, data augmentation, regularized fine-tuning, and more. Despite the development, it is nontrivial to reach a principled understanding of fairness and an effective algorithm that can consistently debias language models. In this work, by rigorous evaluations of Neural Collapse -- a learning phenomenon happen in last-layer representations and classifiers in deep networks -- on fairness-related words, we find that debiased language models exhibit collapsed alignment between token representations and word embeddings. More importantly, this observation inspires us to design a principled fine-tuning method that can effectively improve fairness in a wide range of debiasing methods, while still preserving the performance of language models on standard natural language understanding tasks. We attach our code at https://github.com/Xujxyang/Fairness-NC-main.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04472
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Collapsed Language Models Promote Fairness
Xu, Jingxuan
Chen, Wuyang
Li, Linyi
Zhao, Yao
Wei, Yunchao
Computation and Language
Computers and Society
To mitigate societal biases implicitly encoded in recent successful pretrained language models, a diverse array of approaches have been proposed to encourage model fairness, focusing on prompting, data augmentation, regularized fine-tuning, and more. Despite the development, it is nontrivial to reach a principled understanding of fairness and an effective algorithm that can consistently debias language models. In this work, by rigorous evaluations of Neural Collapse -- a learning phenomenon happen in last-layer representations and classifiers in deep networks -- on fairness-related words, we find that debiased language models exhibit collapsed alignment between token representations and word embeddings. More importantly, this observation inspires us to design a principled fine-tuning method that can effectively improve fairness in a wide range of debiasing methods, while still preserving the performance of language models on standard natural language understanding tasks. We attach our code at https://github.com/Xujxyang/Fairness-NC-main.
title Collapsed Language Models Promote Fairness
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2410.04472