Unsupervised Concept Vector Extraction for Bias Control in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cyberey, Hannah, Ji, Yangfeng, Evans, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912590905475072
author Cyberey, Hannah
Ji, Yangfeng
Evans, David
author_facet Cyberey, Hannah
Ji, Yangfeng
Evans, David
contents Large language models (LLMs) are known to perpetuate stereotypes and exhibit biases. Various strategies have been proposed to mitigate these biases, but most work studies biases as a black-box problem without considering how concepts are represented within the model. We adapt techniques from representation engineering to study how the concept of "gender" is represented within LLMs. We introduce a new method that extracts concept representations via probability weighting without labeled data and efficiently selects a steering vector for measuring and manipulating the model's representation. We develop a projection-based method that enables precise steering of model predictions and demonstrate its effectiveness in mitigating gender bias in LLMs and show that it also generalizes to racial bias. Our code is available at: https://github.com/hannahxchen/gender-bias-steering
format Preprint
id arxiv_https___arxiv_org_abs_2502_19721
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Concept Vector Extraction for Bias Control in LLMs
Cyberey, Hannah
Ji, Yangfeng
Evans, David
Computation and Language
Computers and Society
Large language models (LLMs) are known to perpetuate stereotypes and exhibit biases. Various strategies have been proposed to mitigate these biases, but most work studies biases as a black-box problem without considering how concepts are represented within the model. We adapt techniques from representation engineering to study how the concept of "gender" is represented within LLMs. We introduce a new method that extracts concept representations via probability weighting without labeled data and efficiently selects a steering vector for measuring and manipulating the model's representation. We develop a projection-based method that enables precise steering of model predictions and demonstrate its effectiveness in mitigating gender bias in LLMs and show that it also generalizes to racial bias. Our code is available at: https://github.com/hannahxchen/gender-bias-steering
title Unsupervised Concept Vector Extraction for Bias Control in LLMs
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2502.19721