To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hedström, Anna, Amoukou, Salim I., Bewley, Tom, Mishra, Saumitra, Veloso, Manuela
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908594323062784
author Hedström, Anna
Amoukou, Salim I.
Bewley, Tom
Mishra, Saumitra
Veloso, Manuela
author_facet Hedström, Anna
Amoukou, Salim I.
Bewley, Tom
Mishra, Saumitra
Veloso, Manuela
contents We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely on fixed, manually tuned steering strengths, often resulting in under or oversteering, MERA addresses these limitations by (i) optimising the intervention direction, and (ii) calibrating when, and how much to steer, thereby provably improving performance or abstaining when no confident correction is possible. Experiments across diverse datasets, and LM families demonstrate safe, effective, non-degrading error correction, and that MERA outperforms existing baselines. Moreover, MERA can be applied on top of existing steering techniques to further enhance their performance, establishing it as a general-purpose, and efficient approach to mechanistic activation steering.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13290
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
Hedström, Anna
Amoukou, Salim I.
Bewley, Tom
Mishra, Saumitra
Veloso, Manuela
Machine Learning
Artificial Intelligence
We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely on fixed, manually tuned steering strengths, often resulting in under or oversteering, MERA addresses these limitations by (i) optimising the intervention direction, and (ii) calibrating when, and how much to steer, thereby provably improving performance or abstaining when no confident correction is possible. Experiments across diverse datasets, and LM families demonstrate safe, effective, non-degrading error correction, and that MERA outperforms existing baselines. Moreover, MERA can be applied on top of existing steering techniques to further enhance their performance, establishing it as a general-purpose, and efficient approach to mechanistic activation steering.
title To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.13290