A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915757431980032 |
|---|---|
| author | O'Brien, Claire Seto, Jessica Roy, Dristi Dwivedi, Aditya Dev, Sunishchal Zhu, Kevin O'Brien, Sean Panda, Ashwinee Lagasse, Ryan |
| author_facet | O'Brien, Claire Seto, Jessica Roy, Dristi Dwivedi, Aditya Dev, Sunishchal Zhu, Kevin O'Brien, Sean Panda, Ashwinee Lagasse, Ryan |
| contents | Behavioral alignment in large language models (LLMs) is often achieved through broad fine-tuning, which can result in undesired side effects like distributional shift and low interpretability. We propose a method for alignment that identifies and updates only the neurons most responsible for a given behavior, a targeted approach that allows for fine-tuning with significantly less data. Using sparse autoencoders (SAEs) and linear probes, we isolate the 3% of MLP neurons most predictive of a target behavior, decode them into residual space, and fine-tune only those neurons using gradient masking. We demonstrate this approach on the task of reducing sycophantic behavior, where our method matches or exceeds state-of-the-art performance on four benchmarks (Syco-Bench, NLP, POLI, PHIL) using Gemma-2-2B and 9B models. Our results show that sparse, neuron-level updates offer a scalable and precise alternative to full-model fine-tuning, remaining effective even in situations when little data is available |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_18939 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy O'Brien, Claire Seto, Jessica Roy, Dristi Dwivedi, Aditya Dev, Sunishchal Zhu, Kevin O'Brien, Sean Panda, Ashwinee Lagasse, Ryan Machine Learning Behavioral alignment in large language models (LLMs) is often achieved through broad fine-tuning, which can result in undesired side effects like distributional shift and low interpretability. We propose a method for alignment that identifies and updates only the neurons most responsible for a given behavior, a targeted approach that allows for fine-tuning with significantly less data. Using sparse autoencoders (SAEs) and linear probes, we isolate the 3% of MLP neurons most predictive of a target behavior, decode them into residual space, and fine-tune only those neurons using gradient masking. We demonstrate this approach on the task of reducing sycophantic behavior, where our method matches or exceeds state-of-the-art performance on four benchmarks (Syco-Bench, NLP, POLI, PHIL) using Gemma-2-2B and 9B models. Our results show that sparse, neuron-level updates offer a scalable and precise alternative to full-model fine-tuning, remaining effective even in situations when little data is available |
| title | A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2601.18939 |