Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Haider, Muhammad Umair, Rizwan, Hammad, Sajjad, Hassan, Ju, Peizhong, Siddique, A. B.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915930008715264
author Haider, Muhammad Umair
Rizwan, Hammad
Sajjad, Hassan
Ju, Peizhong
Siddique, A. B.
author_facet Haider, Muhammad Umair
Rizwan, Hammad
Sajjad, Hassan
Ju, Peizhong
Siddique, A. B.
contents Pervasive polysemanticity in large language models (LLMs) undermines discrete neuron-concept attribution, posing a significant challenge for model interpretation and control. We systematically analyze both encoder and decoder based LLMs across diverse datasets, and observe that even highly salient neurons for specific semantic concepts consistently exhibit polysemantic behavior. Importantly, we uncover a consistent pattern: concept-conditioned activation magnitudes of neurons form distinct, often Gaussian-like distributions with minimal overlap. Building on this observation, we hypothesize that interpreting and intervening on concept-specific activation ranges can enable more precise interpretability and targeted manipulation in LLMs. To this end, we introduce NeuronLens, a novel range-based interpretation and manipulation framework, that localizes concept attribution to activation ranges within a neuron. Extensive empirical evaluations show that range-based interventions enable effective manipulation of target concepts while causing substantially less collateral degradation to auxiliary concepts and overall model performance compared to neuron-level masking.
format Preprint
id arxiv_https___arxiv_org_abs_2502_06809
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution
Haider, Muhammad Umair
Rizwan, Hammad
Sajjad, Hassan
Ju, Peizhong
Siddique, A. B.
Machine Learning
Artificial Intelligence
Computation and Language
Pervasive polysemanticity in large language models (LLMs) undermines discrete neuron-concept attribution, posing a significant challenge for model interpretation and control. We systematically analyze both encoder and decoder based LLMs across diverse datasets, and observe that even highly salient neurons for specific semantic concepts consistently exhibit polysemantic behavior. Importantly, we uncover a consistent pattern: concept-conditioned activation magnitudes of neurons form distinct, often Gaussian-like distributions with minimal overlap. Building on this observation, we hypothesize that interpreting and intervening on concept-specific activation ranges can enable more precise interpretability and targeted manipulation in LLMs. To this end, we introduce NeuronLens, a novel range-based interpretation and manipulation framework, that localizes concept attribution to activation ranges within a neuron. Extensive empirical evaluations show that range-based interventions enable effective manipulation of target concepts while causing substantially less collateral degradation to auxiliary concepts and overall model performance compared to neuron-level masking.
title Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.06809