Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911461430788096 |
|---|---|
| author | Cho, Seong Hah Li, Junyi Leshinskaya, Anna |
| author_facet | Cho, Seong Hah Li, Junyi Leshinskaya, Anna |
| contents | Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Among the characteristics of value representation in humans is that they distinguish among value of different kinds. We investigate whether LLMs likewise distinguish three different kinds of good: moral, grammatical, and economic. By probing model behavior, embeddings, and residual stream activations, we report pervasive cases of value entanglement: a conflation between these distinct representations of value. Specifically, both grammatical and economic valuation was found to be overly influenced by moral value, relative to human norms. This conflation was repaired by selective ablation of the activation vectors associated with morality. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_19101 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Cho, Seong Hah Li, Junyi Leshinskaya, Anna Computation and Language Artificial Intelligence Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Among the characteristics of value representation in humans is that they distinguish among value of different kinds. We investigate whether LLMs likewise distinguish three different kinds of good: moral, grammatical, and economic. By probing model behavior, embeddings, and residual stream activations, we report pervasive cases of value entanglement: a conflation between these distinct representations of value. Specifically, both grammatical and economic valuation was found to be overly influenced by moral value, relative to human norms. This conflation was repaired by selective ablation of the activation vectors associated with morality. |
| title | Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2602.19101 |