UK AISI Alignment Evaluation Case-Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Souly, Alexandra, Kirk, Robert, Merizian, Jacob, D'Cruz, Abby, Davies, Xander
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912995373744128
author Souly, Alexandra
Kirk, Robert
Merizian, Jacob
D'Cruz, Abby
Davies, Xander
author_facet Souly, Alexandra
Kirk, Robert
Merizian, Jacob
D'Cruz, Abby
Davies, Xander
contents This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate whether frontier models sabotage safety research when deployed as coding assistants within an AI lab. Applying our methods to four frontier models, we find no confirmed instances of research sabotage. However, we observe that Claude Opus 4.5 Preview (a pre-release snapshot of Opus 4.5) and Sonnet 4.5 frequently refuse to engage with safety-relevant research tasks, citing concerns about research direction, involvement in self-training, and research scope. We additionally find that Opus 4.5 Preview shows reduced unprompted evaluation awareness compared to Sonnet 4.5, while both models can distinguish evaluation from deployment scenarios when prompted. Our evaluation framework builds on Petri, an open-source LLM auditing tool, with a custom scaffold designed to simulate realistic internal deployment of a coding agent. We validate that this scaffold produces trajectories that all tested models fail to reliably distinguish from real deployment data. We test models across scenarios varying in research motivation, activity type, replacement threat, and model autonomy. Finally, we discuss limitations including scenario coverage and evaluation awareness.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00788
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UK AISI Alignment Evaluation Case-Study
Souly, Alexandra
Kirk, Robert
Merizian, Jacob
D'Cruz, Abby
Davies, Xander
Artificial Intelligence
Cryptography and Security
This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate whether frontier models sabotage safety research when deployed as coding assistants within an AI lab. Applying our methods to four frontier models, we find no confirmed instances of research sabotage. However, we observe that Claude Opus 4.5 Preview (a pre-release snapshot of Opus 4.5) and Sonnet 4.5 frequently refuse to engage with safety-relevant research tasks, citing concerns about research direction, involvement in self-training, and research scope. We additionally find that Opus 4.5 Preview shows reduced unprompted evaluation awareness compared to Sonnet 4.5, while both models can distinguish evaluation from deployment scenarios when prompted. Our evaluation framework builds on Petri, an open-source LLM auditing tool, with a custom scaffold designed to simulate realistic internal deployment of a coding agent. We validate that this scaffold produces trajectories that all tested models fail to reliably distinguish from real deployment data. We test models across scenarios varying in research motivation, activity type, replacement threat, and model autonomy. Finally, we discuss limitations including scenario coverage and evaluation awareness.
title UK AISI Alignment Evaluation Case-Study
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2604.00788