Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Raimondi, Bianca, Dalbagno, Daniela, Gabbrielli, Maurizio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914181408620544
author Raimondi, Bianca
Dalbagno, Daniela
Gabbrielli, Maurizio
author_facet Raimondi, Bianca
Dalbagno, Daniela
Gabbrielli, Maurizio
contents Large language models (LLMs) have been shown to internalize human-like biases during finetuning, yet the mechanisms by which these biases manifest remain unclear. In this work, we investigated whether the well-known Knobe effect, a moral bias in intentionality judgements, emerges in finetuned LLMs and whether it can be traced back to specific components of the model. We conducted a Layer-Patching analysis across 3 open-weights LLMs and demonstrated that the bias is not only learned during finetuning but also localized in a specific set of layers. Surprisingly, we found that patching activations from the corresponding pretrained model into just a few critical layers is sufficient to eliminate the effect. Our findings offer new evidence that social biases in LLMs can be interpreted, localized, and mitigated through targeted interventions, without the need for model retraining.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12229
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
Raimondi, Bianca
Dalbagno, Daniela
Gabbrielli, Maurizio
Computation and Language
Artificial Intelligence
Large language models (LLMs) have been shown to internalize human-like biases during finetuning, yet the mechanisms by which these biases manifest remain unclear. In this work, we investigated whether the well-known Knobe effect, a moral bias in intentionality judgements, emerges in finetuned LLMs and whether it can be traced back to specific components of the model. We conducted a Layer-Patching analysis across 3 open-weights LLMs and demonstrated that the bias is not only learned during finetuning but also localized in a specific set of layers. Surprisingly, we found that patching activations from the corresponding pretrained model into just a few critical layers is sufficient to eliminate the effect. Our findings offer new evidence that social biases in LLMs can be interpreted, localized, and mitigated through targeted interventions, without the need for model retraining.
title Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.12229