Emergent Introspection in AI is Content-Agnostic

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lederman, Harvey, Mahowald, Kyle
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913011320487936
author Lederman, Harvey
Mahowald, Kyle
author_facet Lederman, Harvey
Mahowald, Kyle
contents Introspection is a foundational cognitive ability, but its mechanism is not well understood. Recent work has shown that AI models can introspect. We study the mechanism of this introspection. We first extensively replicate Lindsey (2025)'s thought injection detection paradigm in large open-source models. We show that introspection in these models is content-agnostic: models can detect that an anomaly occurred even when they cannot reliably identify its content. The models confabulate injected concepts that are high-frequency and concrete (e.g., "apple"). They also require fewer tokens to detect an injection than to guess the correct concept (with wrong guesses coming earlier). We argue that a content-agnostic introspective mechanism is consistent with leading theories in philosophy and psychology.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05414
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Emergent Introspection in AI is Content-Agnostic
Lederman, Harvey
Mahowald, Kyle
Artificial Intelligence
Computation and Language
Introspection is a foundational cognitive ability, but its mechanism is not well understood. Recent work has shown that AI models can introspect. We study the mechanism of this introspection. We first extensively replicate Lindsey (2025)'s thought injection detection paradigm in large open-source models. We show that introspection in these models is content-agnostic: models can detect that an anomaly occurred even when they cannot reliably identify its content. The models confabulate injected concepts that are high-frequency and concrete (e.g., "apple"). They also require fewer tokens to detect an injection than to guess the correct concept (with wrong guesses coming earlier). We argue that a content-agnostic introspective mechanism is consistent with leading theories in philosophy and psychology.
title Emergent Introspection in AI is Content-Agnostic
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.05414