No Training Wheels: Steering Vectors for Bias Correction at Inference Time

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Aviral, Sethi, Armaan, Sethi, Ameesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912444684697600
author Gupta, Aviral
Sethi, Armaan
Sethi, Ameesh
author_facet Gupta, Aviral
Sethi, Armaan
Sethi, Ameesh
contents Neural network classifiers trained on datasets with uneven group representation often inherit class biases and learn spurious correlations. These models may perform well on average but consistently fail on atypical groups. For example, in hair color classification, datasets may over-represent females with blond hair, reinforcing stereotypes. Although various algorithmic and data-centric methods have been proposed to address such biases, they often require retraining or significant compute. In this work, we propose a cheap, training-free method inspired by steering vectors used to edit behaviors in large language models. We compute the difference in mean activations between majority and minority groups to define a "bias vector," which we subtract from the model's residual stream. This leads to reduced classification bias and improved worst-group accuracy. We explore multiple strategies for extracting and applying these vectors in transformer-like classifiers, showing that steering vectors, traditionally used in generative models, can also be effective in classification. More broadly, we showcase an extremely cheap, inference time, training free method to mitigate bias in classification models.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18598
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle No Training Wheels: Steering Vectors for Bias Correction at Inference Time
Gupta, Aviral
Sethi, Armaan
Sethi, Ameesh
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Neural network classifiers trained on datasets with uneven group representation often inherit class biases and learn spurious correlations. These models may perform well on average but consistently fail on atypical groups. For example, in hair color classification, datasets may over-represent females with blond hair, reinforcing stereotypes. Although various algorithmic and data-centric methods have been proposed to address such biases, they often require retraining or significant compute. In this work, we propose a cheap, training-free method inspired by steering vectors used to edit behaviors in large language models. We compute the difference in mean activations between majority and minority groups to define a "bias vector," which we subtract from the model's residual stream. This leads to reduced classification bias and improved worst-group accuracy. We explore multiple strategies for extracting and applying these vectors in transformer-like classifiers, showing that steering vectors, traditionally used in generative models, can also be effective in classification. More broadly, we showcase an extremely cheap, inference time, training free method to mitigate bias in classification models.
title No Training Wheels: Steering Vectors for Bias Correction at Inference Time
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.18598