Learning to Steer: Input-dependent Steering for Multimodal LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Parekh, Jayneel, Khayatan, Pegah, Shukor, Mustafa, Dapogny, Arnaud, Newson, Alasdair, Cord, Matthieu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908623998812160
author Parekh, Jayneel
Khayatan, Pegah
Shukor, Mustafa
Dapogny, Arnaud
Newson, Alasdair
Cord, Matthieu
author_facet Parekh, Jayneel
Khayatan, Pegah
Shukor, Mustafa
Dapogny, Arnaud
Newson, Alasdair
Cord, Matthieu
contents Steering has emerged as a practical approach to enable post-hoc guidance of LLMs towards enforcing a specific behavior. However, it remains largely underexplored for multimodal LLMs (MLLMs); furthermore, existing steering techniques, such as mean steering, rely on a single steering vector, applied independently of the input query. This paradigm faces limitations when the desired behavior is dependent on the example at hand. For example, a safe answer may consist in abstaining from answering when asked for an illegal activity, or may point to external resources or consultation with an expert when asked about medical advice. In this paper, we investigate a fine-grained steering that uses an input-specific linear shift. This shift is computed using contrastive input-specific prompting. However, the input-specific prompts required for this approach are not known at test time. Therefore, we propose to train a small auxiliary module to predict the input-specific steering vector. Our approach, dubbed as L2S (Learn-to-Steer), demonstrates that it reduces hallucinations and enforces safety in MLLMs, outperforming other static baselines. Our code is publicly available at https://jayneelparekh.github.io/learn-to-steer/
format Preprint
id arxiv_https___arxiv_org_abs_2508_12815
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Steer: Input-dependent Steering for Multimodal LLMs
Parekh, Jayneel
Khayatan, Pegah
Shukor, Mustafa
Dapogny, Arnaud
Newson, Alasdair
Cord, Matthieu
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Steering has emerged as a practical approach to enable post-hoc guidance of LLMs towards enforcing a specific behavior. However, it remains largely underexplored for multimodal LLMs (MLLMs); furthermore, existing steering techniques, such as mean steering, rely on a single steering vector, applied independently of the input query. This paradigm faces limitations when the desired behavior is dependent on the example at hand. For example, a safe answer may consist in abstaining from answering when asked for an illegal activity, or may point to external resources or consultation with an expert when asked about medical advice. In this paper, we investigate a fine-grained steering that uses an input-specific linear shift. This shift is computed using contrastive input-specific prompting. However, the input-specific prompts required for this approach are not known at test time. Therefore, we propose to train a small auxiliary module to predict the input-specific steering vector. Our approach, dubbed as L2S (Learn-to-Steer), demonstrates that it reduces hallucinations and enforces safety in MLLMs, outperforming other static baselines. Our code is publicly available at https://jayneelparekh.github.io/learn-to-steer/
title Learning to Steer: Input-dependent Steering for Multimodal LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.12815