Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Roytburg, Dani, Bozoukov, Matthew, Nguyen, Matthew, Barzdukas, Jou, Fu, Simon, Oozeer, Narmeen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918135267852288
author Roytburg, Dani
Bozoukov, Matthew
Nguyen, Matthew
Barzdukas, Jou
Fu, Simon
Oozeer, Narmeen
author_facet Roytburg, Dani
Bozoukov, Matthew
Nguyen, Matthew
Barzdukas, Jou
Fu, Simon
Oozeer, Narmeen
contents Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models. This bias undermines fairness and reliability in evaluation pipelines, particularly for tasks like preference tuning and model routing. We investigate whether lightweight steering vectors can mitigate this problem at inference time without retraining. We introduce a curated dataset that distinguishes self-preference bias into justified examples of self-preference and unjustified examples of self-preference, and we construct steering vectors using two methods: Contrastive Activation Addition (CAA) and an optimization-based approach. Our results show that steering vectors can reduce unjustified self-preference bias by up to 97\%, substantially outperforming prompting and direct preference optimization baselines. Yet steering vectors are unstable on legitimate self-preference and unbiased agreement, implying self-preference spans multiple or nonlinear directions. This underscores both their promise and limits as safeguards for LLM-as-judges and motivates more robust interventions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03647
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
Roytburg, Dani
Bozoukov, Matthew
Nguyen, Matthew
Barzdukas, Jou
Fu, Simon
Oozeer, Narmeen
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models. This bias undermines fairness and reliability in evaluation pipelines, particularly for tasks like preference tuning and model routing. We investigate whether lightweight steering vectors can mitigate this problem at inference time without retraining. We introduce a curated dataset that distinguishes self-preference bias into justified examples of self-preference and unjustified examples of self-preference, and we construct steering vectors using two methods: Contrastive Activation Addition (CAA) and an optimization-based approach. Our results show that steering vectors can reduce unjustified self-preference bias by up to 97\%, substantially outperforming prompting and direct preference optimization baselines. Yet steering vectors are unstable on legitimate self-preference and unbiased agreement, implying self-preference spans multiple or nonlinear directions. This underscores both their promise and limits as safeguards for LLM-as-judges and motivates more robust interventions.
title Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.03647