Safety by Inseparability: Toward Architectures Where Alignment Cannot Be Removed

Fuente: Zenodo
Enregistré dans:
Détails bibliographiques
Auteur principal: Sean Everett, Morin
Format: Recurso digital
Langue:anglais
Publié: Zenodo 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866901726734319616
author Sean Everett, Morin
author_facet Sean Everett, Morin
contents <p>This paper introduces the principle of Safety by Inseparability: an architectural property where the mechanisms responsible for safe behavior are structurally identical to those responsible for reasoning capability. In such architectures, removing safety simultaneously destroys the model's ability to reason, eliminating the possibility   of a "capable but unsafe" configuration. The principle is grounded in multi-expert deliberative architectures with internal quality assessment. Theoretical analysis and four falsifiable predictions are presented. Implementation details are withheld.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19564496
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Safety by Inseparability: Toward Architectures Where Alignment Cannot Be Removed
Sean Everett, Morin
AI safety
alignment
mixture of experts
inseparability
architectural safety
<p>This paper introduces the principle of Safety by Inseparability: an architectural property where the mechanisms responsible for safe behavior are structurally identical to those responsible for reasoning capability. In such architectures, removing safety simultaneously destroys the model's ability to reason, eliminating the possibility   of a "capable but unsafe" configuration. The principle is grounded in multi-expert deliberative architectures with internal quality assessment. Theoretical analysis and four falsifiable predictions are presented. Implementation details are withheld.</p>
title Safety by Inseparability: Toward Architectures Where Alignment Cannot Be Removed
topic AI safety
alignment
mixture of experts
inseparability
architectural safety
url https://doi.org/10.5281/zenodo.19564496