Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahsan, Hiba, Sharma, Arnab Sen, Amir, Silvio, Bau, David, Wallace, Byron C.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918148108713984
author Ahsan, Hiba
Sharma, Arnab Sen
Amir, Silvio
Bau, David
Wallace, Byron C.
author_facet Ahsan, Hiba
Sharma, Arnab Sen
Amir, Silvio
Bau, David
Wallace, Byron C.
contents We know from prior work that LLMs encode social biases, and that this manifests in clinical tasks. In this work we adopt tools from mechanistic interpretability to unveil sociodemographic representations and biases within LLMs in the context of healthcare. Specifically, we ask: Can we identify activations within LLMs that encode sociodemographic information (e.g., gender, race)? We find that gender information is highly localized in MLP layers and can be reliably manipulated at inference time via patching. Such interventions can surgically alter generated clinical vignettes for specific conditions, and also influence downstream clinical predictions which correlate with gender, e.g., patient risk of depression. We find that representation of patient race is somewhat more distributed, but can also be intervened upon, to a degree. To our knowledge, this is the first application of mechanistic interpretability methods to LLMs for healthcare.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13319
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
Ahsan, Hiba
Sharma, Arnab Sen
Amir, Silvio
Bau, David
Wallace, Byron C.
Computation and Language
We know from prior work that LLMs encode social biases, and that this manifests in clinical tasks. In this work we adopt tools from mechanistic interpretability to unveil sociodemographic representations and biases within LLMs in the context of healthcare. Specifically, we ask: Can we identify activations within LLMs that encode sociodemographic information (e.g., gender, race)? We find that gender information is highly localized in MLP layers and can be reliably manipulated at inference time via patching. Such interventions can surgically alter generated clinical vignettes for specific conditions, and also influence downstream clinical predictions which correlate with gender, e.g., patient risk of depression. We find that representation of patient race is somewhat more distributed, but can also be intervened upon, to a degree. To our knowledge, this is the first application of mechanistic interpretability methods to LLMs for healthcare.
title Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
topic Computation and Language
url https://arxiv.org/abs/2502.13319