Conceptors for Semantic Steering
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917464356421632 |
|---|---|
| author | Triantafyllopoulos, Ilias Cho, Young-Min Tao, Ren Miao, Miranda Muqing Rai, Sunny Ungar, Lyle Guntuku, Sharath Chandra Ryant, Neville Sedoc, João |
| author_facet | Triantafyllopoulos, Ilias Cho, Young-Min Tao, Ren Miao, Miranda Muqing Rai, Sunny Ungar, Lyle Guntuku, Sharath Chandra Ryant, Neville Sedoc, João |
| contents | Activation-based steering provides control of LLM behavior at inference time, but the dominant paradigm reduces each concept to a single direction whose geometry is left largely unexamined. Rather than selecting a single steering direction, we use conceptors: soft projection matrices estimated from activations pooled across both poles of a bipolar concept, which preserve the concept's full multidimensional subspace. A geometric analysis shows the bipolar subspace strictly subsumes the single-vector baseline. We further show that the conceptor quota provides a parameter-free layer-selection diagnostic, predicting concept separability with Pearson correlations up to r=0.96 across three instruction-tuned models and three semantic dimensions. Beyond selection, conceptors admit a closed-form Boolean algebra (AND, OR, NOT): we evaluate conceptor compositionality on thematically related sub-concepts. Across a systematic five-axis design-space evaluation, conceptors match or outperform additive baselines at layers where concept subspaces are multi-dimensional while producing substantially fewer degenerate outputs. Conceptor steering is a geometrically principled, compositional, and practically safer alternative to single-direction steering from a limited number of contrastive pairs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_04980 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Conceptors for Semantic Steering Triantafyllopoulos, Ilias Cho, Young-Min Tao, Ren Miao, Miranda Muqing Rai, Sunny Ungar, Lyle Guntuku, Sharath Chandra Ryant, Neville Sedoc, João Machine Learning Computation and Language Activation-based steering provides control of LLM behavior at inference time, but the dominant paradigm reduces each concept to a single direction whose geometry is left largely unexamined. Rather than selecting a single steering direction, we use conceptors: soft projection matrices estimated from activations pooled across both poles of a bipolar concept, which preserve the concept's full multidimensional subspace. A geometric analysis shows the bipolar subspace strictly subsumes the single-vector baseline. We further show that the conceptor quota provides a parameter-free layer-selection diagnostic, predicting concept separability with Pearson correlations up to r=0.96 across three instruction-tuned models and three semantic dimensions. Beyond selection, conceptors admit a closed-form Boolean algebra (AND, OR, NOT): we evaluate conceptor compositionality on thematically related sub-concepts. Across a systematic five-axis design-space evaluation, conceptors match or outperform additive baselines at layers where concept subspaces are multi-dimensional while producing substantially fewer degenerate outputs. Conceptor steering is a geometrically principled, compositional, and practically safer alternative to single-direction steering from a limited number of contrastive pairs. |
| title | Conceptors for Semantic Steering |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2605.04980 |