Emergent Self-Monitoring in Large Language Models: Probing Internal State Awareness and Output Ownership

Fuente: Zenodo
Enregistré dans:
Détails bibliographiques
Auteur principal: Nadeem, Aurther
Format: Recurso digital
Langue:anglais
Publié: Zenodo 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866902236843474944
author Nadeem, Aurther
author_facet Nadeem, Aurther
contents <p>We study whether large language models can monitor and report on perturbations to their own internal activations, rather than merely role-playing about their “thoughts.” Using latent-space interventions and activation-level measurements, we design four experiments probing detection of injected concepts, attribution of thought origin, ownership of outputs under forced generation, and volitional latent steering.</p> <p>Evaluating Llama 3.1 and 3.3 models across scales and tuning variants, we find a consistent dissociation between internal control and introspective awareness: models can sometimes steer latent state deliberately, yet fail to reliably detect, attribute, or take ownership of externally induced internal changes. Instruction tuning reshapes introspective policy without improving discrimination.</p> <p>Code and datasets are publicly released to support further benchmarking of internal-state awareness in language models.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18027539
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Emergent Self-Monitoring in Large Language Models: Probing Internal State Awareness and Output Ownership
Nadeem, Aurther
large language models
interpretability
mechanistic interpretability
introspection
activation engineering
alignment
<p>We study whether large language models can monitor and report on perturbations to their own internal activations, rather than merely role-playing about their “thoughts.” Using latent-space interventions and activation-level measurements, we design four experiments probing detection of injected concepts, attribution of thought origin, ownership of outputs under forced generation, and volitional latent steering.</p> <p>Evaluating Llama 3.1 and 3.3 models across scales and tuning variants, we find a consistent dissociation between internal control and introspective awareness: models can sometimes steer latent state deliberately, yet fail to reliably detect, attribute, or take ownership of externally induced internal changes. Instruction tuning reshapes introspective policy without improving discrimination.</p> <p>Code and datasets are publicly released to support further benchmarking of internal-state awareness in language models.</p>
title Emergent Self-Monitoring in Large Language Models: Probing Internal State Awareness and Output Ownership
topic large language models
interpretability
mechanistic interpretability
introspection
activation engineering
alignment
url https://doi.org/10.5281/zenodo.18027539