Focus On This, Not That! Steering LLMs with Adaptive Feature Specification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lamb, Tom A., Davies, Adam, Paren, Alasdair, Torr, Philip H. S., Pinto, Francesco
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912413595467776
author Lamb, Tom A.
Davies, Adam
Paren, Alasdair
Torr, Philip H. S.
Pinto, Francesco
author_facet Lamb, Tom A.
Davies, Adam
Paren, Alasdair
Torr, Philip H. S.
Pinto, Francesco
contents Despite the success of Instruction Tuning (IT) in training large language models (LLMs), such models often leverage spurious or biased features learnt from their training data and can become misaligned, leading to undesired behaviours. While existing techniques can steer model behaviour at inference-time, they are often post-hoc and do not embed steering as an intrinsic model feature. In this work, we introduce Focus Instruction Tuning (FIT), which trains LLMs to condition their responses by focusing on specific features whilst ignoring others, leading to different behaviours based on what features are specified. Across diverse benchmarks, we demonstrate that FIT: (i) successfully steers behaviour at inference time; (ii) increases robustness by amplifying core task signals and down-weighting spurious cues; (iii) mitigates social bias by suppressing demographic attributes; and (iv) generalises under distribution shifts and to previously unseen focus features. FIT therefore offers a lightweight, intrinsic mechanism for building more robust, fair, and easily controllable LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22944
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
Lamb, Tom A.
Davies, Adam
Paren, Alasdair
Torr, Philip H. S.
Pinto, Francesco
Machine Learning
Artificial Intelligence
Computation and Language
Despite the success of Instruction Tuning (IT) in training large language models (LLMs), such models often leverage spurious or biased features learnt from their training data and can become misaligned, leading to undesired behaviours. While existing techniques can steer model behaviour at inference-time, they are often post-hoc and do not embed steering as an intrinsic model feature. In this work, we introduce Focus Instruction Tuning (FIT), which trains LLMs to condition their responses by focusing on specific features whilst ignoring others, leading to different behaviours based on what features are specified. Across diverse benchmarks, we demonstrate that FIT: (i) successfully steers behaviour at inference time; (ii) increases robustness by amplifying core task signals and down-weighting spurious cues; (iii) mitigates social bias by suppressing demographic attributes; and (iv) generalises under distribution shifts and to previously unseen focus features. FIT therefore offers a lightweight, intrinsic mechanism for building more robust, fair, and easily controllable LLMs.
title Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.22944