MidSteer: Optimal Affine Framework for Steering Generative Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gaintseva, Tatiana, Stepanov, Andrew, Liu, Ziquan, Benning, Martin, Slabaugh, Gregory, Deng, Jiankang, Elezi, Ismail
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918532114022400
author Gaintseva, Tatiana
Stepanov, Andrew
Liu, Ziquan
Benning, Martin
Slabaugh, Gregory
Deng, Jiankang
Elezi, Ismail
author_facet Gaintseva, Tatiana
Stepanov, Andrew
Liu, Ziquan
Benning, Martin
Slabaugh, Gregory
Deng, Jiankang
Elezi, Ismail
contents Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment alignment and safety settings. However, despite its empirical success, it currently lacks a comprehensive theoretical framework. In this paper, we bridge this gap by formalizing the theory of concept steering. First, we establish a link between steering and affine concept erasure, proving that the standard approach for removing unwanted behaviors is a special case of LEACE (a closed-form method for affine erasure). Next, we formulate a principled theoretical framework for concept switching, LEACE-Switch, and characterize the assumptions under which it provides an optimal affine solution. Building on this analysis, we then introduce MidSteer (Minimal Disturbance concept Steering), a more general affine framework for concept manipulation that relaxes these assumptions and enables directed, minimal-disturbance transformations. We demonstrate that MidSteer performs favorably across a range of tasks, modalities, and architectures, including vision diffusion models and large language models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05220
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MidSteer: Optimal Affine Framework for Steering Generative Models
Gaintseva, Tatiana
Stepanov, Andrew
Liu, Ziquan
Benning, Martin
Slabaugh, Gregory
Deng, Jiankang
Elezi, Ismail
Machine Learning
Artificial Intelligence
Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment alignment and safety settings. However, despite its empirical success, it currently lacks a comprehensive theoretical framework. In this paper, we bridge this gap by formalizing the theory of concept steering. First, we establish a link between steering and affine concept erasure, proving that the standard approach for removing unwanted behaviors is a special case of LEACE (a closed-form method for affine erasure). Next, we formulate a principled theoretical framework for concept switching, LEACE-Switch, and characterize the assumptions under which it provides an optimal affine solution. Building on this analysis, we then introduce MidSteer (Minimal Disturbance concept Steering), a more general affine framework for concept manipulation that relaxes these assumptions and enables directed, minimal-disturbance transformations. We demonstrate that MidSteer performs favorably across a range of tasks, modalities, and architectures, including vision diffusion models and large language models.
title MidSteer: Optimal Affine Framework for Steering Generative Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.05220