Augmenting Bias Detection in LLMs Using Topological Data Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Varadarajan, Keshav, Songdechakraiwut, Tananun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915438184628224
author Varadarajan, Keshav
Songdechakraiwut, Tananun
author_facet Varadarajan, Keshav
Songdechakraiwut, Tananun
contents Recently, many bias detection methods have been proposed to determine the level of bias a large language model captures. However, tests to identify which parts of a large language model are responsible for bias towards specific groups remain underdeveloped. In this study, we present a method using topological data analysis to identify which heads in GPT-2 contribute to the misrepresentation of identity groups present in the StereoSet dataset. We find that biases for particular categories, such as gender or profession, are concentrated in attention heads that act as hot spots. The metric we propose can also be used to determine which heads capture bias for a specific group within a bias category, and future work could extend this method to help de-bias large language models.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07516
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Augmenting Bias Detection in LLMs Using Topological Data Analysis
Varadarajan, Keshav
Songdechakraiwut, Tananun
Computation and Language
Recently, many bias detection methods have been proposed to determine the level of bias a large language model captures. However, tests to identify which parts of a large language model are responsible for bias towards specific groups remain underdeveloped. In this study, we present a method using topological data analysis to identify which heads in GPT-2 contribute to the misrepresentation of identity groups present in the StereoSet dataset. We find that biases for particular categories, such as gender or profession, are concentrated in attention heads that act as hot spots. The metric we propose can also be used to determine which heads capture bias for a specific group within a bias category, and future work could extend this method to help de-bias large language models.
title Augmenting Bias Detection in LLMs Using Topological Data Analysis
topic Computation and Language
url https://arxiv.org/abs/2508.07516