Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Weilun, Rusnak, Alexander, Kaplan, Frederic
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:https://arxiv.org/abs/2603.23659
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911543104372736
author Xu, Weilun
Rusnak, Alexander
Kaplan, Frederic
author_facet Xu, Weilun
Rusnak, Alexander
Kaplan, Frederic
contents When large language models make ethical judgments, do their internal representations distinguish between normative frameworks, or collapse ethics into a single acceptability dimension? We probe hidden representations across five ethical frameworks (deontology, utilitarianism, virtue, justice, commonsense) in six LLMs spanning 4B--72B parameters. Our analysis reveals differentiated ethical subspaces with asymmetric transfer patterns -- e.g., deontology probes partially generalize to virtue scenarios while commonsense probes fail catastrophically on justice. Disagreement between deontological and utilitarian probes correlates with higher behavioral entropy across architectures, though this relationship may partly reflect shared sensitivity to scenario difficulty. Post-hoc validation reveals that probes partially depend on surface features of benchmark templates, motivating cautious interpretation. We discuss both the structural insights these methods provide and their epistemological limitations.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23659
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Probing Ethical Framework Representations in Large Language Models: Structure, Entanglement, and Methodological Challenges
Xu, Weilun
Rusnak, Alexander
Kaplan, Frederic
Computation and Language
Artificial Intelligence
When large language models make ethical judgments, do their internal representations distinguish between normative frameworks, or collapse ethics into a single acceptability dimension? We probe hidden representations across five ethical frameworks (deontology, utilitarianism, virtue, justice, commonsense) in six LLMs spanning 4B--72B parameters. Our analysis reveals differentiated ethical subspaces with asymmetric transfer patterns -- e.g., deontology probes partially generalize to virtue scenarios while commonsense probes fail catastrophically on justice. Disagreement between deontological and utilitarian probes correlates with higher behavioral entropy across architectures, though this relationship may partly reflect shared sensitivity to scenario difficulty. Post-hoc validation reveals that probes partially depend on surface features of benchmark templates, motivating cautious interpretation. We discuss both the structural insights these methods provide and their epistemological limitations.
title Probing Ethical Framework Representations in Large Language Models: Structure, Entanglement, and Methodological Challenges
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.23659