Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bao, Yuntai, Zhang, Xuhong, Chen, Jintao, Su, Ge, Cai, Yuxiang, Peng, Hao, Sun, Bing, Weng, Haiqin, Yan, Liu, Yin, Jianwei
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2602.05234
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911517223419904
author Bao, Yuntai
Zhang, Xuhong
Chen, Jintao
Su, Ge
Cai, Yuxiang
Peng, Hao
Sun, Bing
Weng, Haiqin
Yan, Liu
Yin, Jianwei
author_facet Bao, Yuntai
Zhang, Xuhong
Chen, Jintao
Su, Ge
Cai, Yuxiang
Peng, Hao
Sun, Bing
Weng, Haiqin
Yan, Liu
Yin, Jianwei
contents Intervention-based model steering offers a lightweight and interpretable alternative to prompting and fine-tuning. However, by adapting strong optimization objectives from fine-tuning, current methods are susceptible to overfitting and often underperform, sometimes generating unnatural outputs. We hypothesize that this is because effective steering requires the faithful identification of internal model mechanisms, not the enforcement of external preferences. To this end, we build on the principles of distributed alignment search (DAS), the standard for causal variable localization, to propose a new steering method: Concept DAS (CDAS). While we adopt the core mechanism of DAS, distributed interchange intervention (DII), we introduce a novel distribution matching objective tailored for the steering task by aligning intervened output distributions with counterfactual distributions. CDAS differs from prior work in two main ways: first, it learns interventions via weak-supervised distribution matching rather than probability maximization; second, it uses DIIs that naturally enable bi-directional steering and allow steering factors to be derived from data, reducing the effort required for hyperparameter tuning and resulting in more faithful and stable control. On AxBench, a large-scale model steering benchmark, we show that CDAS does not always outperform preference-optimization methods but may benefit more from increased model scale. In two safety-related case studies, overriding refusal behaviors of safety-aligned models and neutralizing a chain-of-thought backdoor, CDAS achieves systematic steering while maintaining general model utility. These results indicate that CDAS is complementary to preference-optimization approaches and conditionally constitutes a robust approach to intervention-based model steering. Our code is available at https://github.com/colored-dye/concept_das.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05234
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions
Bao, Yuntai
Zhang, Xuhong
Chen, Jintao
Su, Ge
Cai, Yuxiang
Peng, Hao
Sun, Bing
Weng, Haiqin
Yan, Liu
Yin, Jianwei
Machine Learning
Computation and Language
Intervention-based model steering offers a lightweight and interpretable alternative to prompting and fine-tuning. However, by adapting strong optimization objectives from fine-tuning, current methods are susceptible to overfitting and often underperform, sometimes generating unnatural outputs. We hypothesize that this is because effective steering requires the faithful identification of internal model mechanisms, not the enforcement of external preferences. To this end, we build on the principles of distributed alignment search (DAS), the standard for causal variable localization, to propose a new steering method: Concept DAS (CDAS). While we adopt the core mechanism of DAS, distributed interchange intervention (DII), we introduce a novel distribution matching objective tailored for the steering task by aligning intervened output distributions with counterfactual distributions. CDAS differs from prior work in two main ways: first, it learns interventions via weak-supervised distribution matching rather than probability maximization; second, it uses DIIs that naturally enable bi-directional steering and allow steering factors to be derived from data, reducing the effort required for hyperparameter tuning and resulting in more faithful and stable control. On AxBench, a large-scale model steering benchmark, we show that CDAS does not always outperform preference-optimization methods but may benefit more from increased model scale. In two safety-related case studies, overriding refusal behaviors of safety-aligned models and neutralizing a chain-of-thought backdoor, CDAS achieves systematic steering while maintaining general model utility. These results indicate that CDAS is complementary to preference-optimization approaches and conditionally constitutes a robust approach to intervention-based model steering. Our code is available at https://github.com/colored-dye/concept_das.
title Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2602.05234