Safety by Inseparability: Toward Architectures Where Alignment Cannot Be Removed
Fuente:
Zenodo
Enregistré dans:
| Auteur principal: | |
|---|---|
| Format: | Recurso digital |
| Langue: | anglais |
| Publié: |
Zenodo
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866901726734319616 |
|---|---|
| author | Sean Everett, Morin |
| author_facet | Sean Everett, Morin |
| contents | <p>This paper introduces the principle of Safety by Inseparability: an architectural property where the mechanisms responsible for safe behavior are structurally identical to those responsible for reasoning capability. In such architectures, removing safety simultaneously destroys the model's ability to reason, eliminating the possibility of a "capable but unsafe" configuration. The principle is grounded in multi-expert deliberative architectures with internal quality assessment. Theoretical analysis and four falsifiable predictions are presented. Implementation details are withheld.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19564496 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Safety by Inseparability: Toward Architectures Where Alignment Cannot Be Removed Sean Everett, Morin AI safety alignment mixture of experts inseparability architectural safety <p>This paper introduces the principle of Safety by Inseparability: an architectural property where the mechanisms responsible for safe behavior are structurally identical to those responsible for reasoning capability. In such architectures, removing safety simultaneously destroys the model's ability to reason, eliminating the possibility of a "capable but unsafe" configuration. The principle is grounded in multi-expert deliberative architectures with internal quality assessment. Theoretical analysis and four falsifiable predictions are presented. Implementation details are withheld.</p> |
| title | Safety by Inseparability: Toward Architectures Where Alignment Cannot Be Removed |
| topic | AI safety alignment mixture of experts inseparability architectural safety |
| url | https://doi.org/10.5281/zenodo.19564496 |