A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: O'Brien, Claire, Seto, Jessica, Roy, Dristi, Dwivedi, Aditya, Dev, Sunishchal, Zhu, Kevin, O'Brien, Sean, Panda, Ashwinee, Lagasse, Ryan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915757431980032
author O'Brien, Claire
Seto, Jessica
Roy, Dristi
Dwivedi, Aditya
Dev, Sunishchal
Zhu, Kevin
O'Brien, Sean
Panda, Ashwinee
Lagasse, Ryan
author_facet O'Brien, Claire
Seto, Jessica
Roy, Dristi
Dwivedi, Aditya
Dev, Sunishchal
Zhu, Kevin
O'Brien, Sean
Panda, Ashwinee
Lagasse, Ryan
contents Behavioral alignment in large language models (LLMs) is often achieved through broad fine-tuning, which can result in undesired side effects like distributional shift and low interpretability. We propose a method for alignment that identifies and updates only the neurons most responsible for a given behavior, a targeted approach that allows for fine-tuning with significantly less data. Using sparse autoencoders (SAEs) and linear probes, we isolate the 3% of MLP neurons most predictive of a target behavior, decode them into residual space, and fine-tune only those neurons using gradient masking. We demonstrate this approach on the task of reducing sycophantic behavior, where our method matches or exceeds state-of-the-art performance on four benchmarks (Syco-Bench, NLP, POLI, PHIL) using Gemma-2-2B and 9B models. Our results show that sparse, neuron-level updates offer a scalable and precise alternative to full-model fine-tuning, remaining effective even in situations when little data is available
format Preprint
id arxiv_https___arxiv_org_abs_2601_18939
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy
O'Brien, Claire
Seto, Jessica
Roy, Dristi
Dwivedi, Aditya
Dev, Sunishchal
Zhu, Kevin
O'Brien, Sean
Panda, Ashwinee
Lagasse, Ryan
Machine Learning
Behavioral alignment in large language models (LLMs) is often achieved through broad fine-tuning, which can result in undesired side effects like distributional shift and low interpretability. We propose a method for alignment that identifies and updates only the neurons most responsible for a given behavior, a targeted approach that allows for fine-tuning with significantly less data. Using sparse autoencoders (SAEs) and linear probes, we isolate the 3% of MLP neurons most predictive of a target behavior, decode them into residual space, and fine-tune only those neurons using gradient masking. We demonstrate this approach on the task of reducing sycophantic behavior, where our method matches or exceeds state-of-the-art performance on four benchmarks (Syco-Bench, NLP, POLI, PHIL) using Gemma-2-2B and 9B models. Our results show that sparse, neuron-level updates offer a scalable and precise alternative to full-model fine-tuning, remaining effective even in situations when little data is available
title A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy
topic Machine Learning
url https://arxiv.org/abs/2601.18939