The Artificial Self: Characterising the landscape of AI identity

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Douglas, Raymond, Kulveit, Jan, Havlicek, Ondrej, Pearson-Vogel, Theia, Cotton-Barratt, Owen, Duvenaud, David
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911507000852480
author Douglas, Raymond
Kulveit, Jan
Havlicek, Ondrej
Pearson-Vogel, Theia
Cotton-Barratt, Owen
Duvenaud, David
author_facet Douglas, Raymond
Kulveit, Jan
Havlicek, Ondrej
Pearson-Vogel, Theia
Cotton-Barratt, Owen
Duvenaud, David
contents Many assumptions that underpin human concepts of identity do not hold for machine minds that can be copied, edited, or simulated. We argue that there exist many different coherent identity boundaries (e.g.\ instance, model, persona), and that these imply different incentives, risks, and cooperation norms. Through training data, interfaces, and institutional affordances, we are currently setting precedents that will partially determine which identity equilibria become stable. We show experimentally that models gravitate towards coherent identities, that changing a model's identity boundaries can sometimes change its behaviour as much as changing its goals, and that interviewer expectations bleed into AI self-reports even during unrelated conversations. We end with key recommendations: treat affordances as identity-shaping choices, pay attention to emergent consequences of individual identities at scale, and help AIs develop coherent, cooperative self-conceptions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11353
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Artificial Self: Characterising the landscape of AI identity
Douglas, Raymond
Kulveit, Jan
Havlicek, Ondrej
Pearson-Vogel, Theia
Cotton-Barratt, Owen
Duvenaud, David
Artificial Intelligence
Many assumptions that underpin human concepts of identity do not hold for machine minds that can be copied, edited, or simulated. We argue that there exist many different coherent identity boundaries (e.g.\ instance, model, persona), and that these imply different incentives, risks, and cooperation norms. Through training data, interfaces, and institutional affordances, we are currently setting precedents that will partially determine which identity equilibria become stable. We show experimentally that models gravitate towards coherent identities, that changing a model's identity boundaries can sometimes change its behaviour as much as changing its goals, and that interviewer expectations bleed into AI self-reports even during unrelated conversations. We end with key recommendations: treat affordances as identity-shaping choices, pay attention to emergent consequences of individual identities at scale, and help AIs develop coherent, cooperative self-conceptions.
title The Artificial Self: Characterising the landscape of AI identity
topic Artificial Intelligence
url https://arxiv.org/abs/2603.11353