Can We Locate and Prevent Stereotypes in LLMs?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: D'Souza, Alex
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913053378871296
author D'Souza, Alex
author_facet D'Souza, Alex
contents Stereotypes in large language models (LLMs) can perpetuate harmful societal biases. Despite the widespread use of models, little is known about where these biases reside in the neural network. This study investigates the internal mechanisms of GPT 2 Small and Llama 3.2 to locate stereotype related activations. We explore two approaches: identifying individual contrastive neuron activations that encode stereotypes, and detecting attention heads that contribute heavily to biased outputs. Our experiments aim to map these "bias fingerprints" and provide initial insights for mitigating stereotypes.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19764
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can We Locate and Prevent Stereotypes in LLMs?
D'Souza, Alex
Computation and Language
Artificial Intelligence
Stereotypes in large language models (LLMs) can perpetuate harmful societal biases. Despite the widespread use of models, little is known about where these biases reside in the neural network. This study investigates the internal mechanisms of GPT 2 Small and Llama 3.2 to locate stereotype related activations. We explore two approaches: identifying individual contrastive neuron activations that encode stereotypes, and detecting attention heads that contribute heavily to biased outputs. Our experiments aim to map these "bias fingerprints" and provide initial insights for mitigating stereotypes.
title Can We Locate and Prevent Stereotypes in LLMs?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.19764