Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kopf, Laura, Feldhus, Nils, Bykov, Kirill, Bommer, Philine Lou, Hedström, Anna, Höhne, Marina M. -C., Eberle, Oliver
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914152175370240
author Kopf, Laura
Feldhus, Nils
Bykov, Kirill
Bommer, Philine Lou
Hedström, Anna
Höhne, Marina M. -C.
Eberle, Oliver
author_facet Kopf, Laura
Feldhus, Nils
Bykov, Kirill
Bommer, Philine Lou
Hedström, Anna
Höhne, Marina M. -C.
Eberle, Oliver
contents Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large language models (LLMs) for natural language processing (NLP), current automated neuron-level feature description methods face two key challenges: limited robustness and the assumption that each neuron encodes a single concept (monosemanticity), despite increasing evidence of polysemanticity. This assumption restricts the expressiveness of feature descriptions and limits their ability to capture the full range of behaviors encoded in model internals. To address this, we introduce Polysemantic FeatuRe Identification and Scoring Method (PRISM), a novel framework specifically designed to capture the complexity of features in LLMs. Unlike approaches that assign a single description per neuron, common in many automated interpretability methods in NLP, PRISM produces more nuanced descriptions that account for both monosemantic and polysemantic behavior. We apply PRISM to LLMs and, through extensive benchmarking against existing methods, demonstrate that our approach produces more accurate and faithful feature descriptions, improving both overall description quality (via a description score) and the ability to capture distinct concepts when polysemanticity is present (via a polysemanticity score).
format Preprint
id arxiv_https___arxiv_org_abs_2506_15538
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
Kopf, Laura
Feldhus, Nils
Bykov, Kirill
Bommer, Philine Lou
Hedström, Anna
Höhne, Marina M. -C.
Eberle, Oliver
Machine Learning
Artificial Intelligence
Computation and Language
Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large language models (LLMs) for natural language processing (NLP), current automated neuron-level feature description methods face two key challenges: limited robustness and the assumption that each neuron encodes a single concept (monosemanticity), despite increasing evidence of polysemanticity. This assumption restricts the expressiveness of feature descriptions and limits their ability to capture the full range of behaviors encoded in model internals. To address this, we introduce Polysemantic FeatuRe Identification and Scoring Method (PRISM), a novel framework specifically designed to capture the complexity of features in LLMs. Unlike approaches that assign a single description per neuron, common in many automated interpretability methods in NLP, PRISM produces more nuanced descriptions that account for both monosemantic and polysemantic behavior. We apply PRISM to LLMs and, through extensive benchmarking against existing methods, demonstrate that our approach produces more accurate and faithful feature descriptions, improving both overall description quality (via a description score) and the ability to capture distinct concepts when polysemanticity is present (via a polysemanticity score).
title Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.15538