A sketch of an AI control safety case

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Korbak, Tomek, Clymer, Joshua, Hilton, Benjamin, Shlegeris, Buck, Irving, Geoffrey
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917904876830720
author Korbak, Tomek
Clymer, Joshua
Hilton, Benjamin
Shlegeris, Buck
Irving, Geoffrey
author_facet Korbak, Tomek
Clymer, Joshua
Hilton, Benjamin
Shlegeris, Buck
Irving, Geoffrey
contents As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how developers could construct a "control safety case", which is a structured argument that models are incapable of subverting control measures in order to cause unacceptable outcomes. As a case study, we sketch an argument that a hypothetical LLM agent deployed internally at an AI company won't exfiltrate sensitive information. The sketch relies on evidence from a "control evaluation,"' where a red team deliberately designs models to exfiltrate data in a proxy for the deployment environment. The safety case then hinges on several claims: (1) the red team adequately elicits model capabilities to exfiltrate data, (2) control measures remain at least as effective in deployment, and (3) developers conservatively extrapolate model performance to predict the probability of data exfiltration in deployment. This safety case sketch is a step toward more concrete arguments that can be used to show that a dangerously capable LLM agent is safe to deploy.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17315
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A sketch of an AI control safety case
Korbak, Tomek
Clymer, Joshua
Hilton, Benjamin
Shlegeris, Buck
Irving, Geoffrey
Artificial Intelligence
Cryptography and Security
Software Engineering
As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how developers could construct a "control safety case", which is a structured argument that models are incapable of subverting control measures in order to cause unacceptable outcomes. As a case study, we sketch an argument that a hypothetical LLM agent deployed internally at an AI company won't exfiltrate sensitive information. The sketch relies on evidence from a "control evaluation,"' where a red team deliberately designs models to exfiltrate data in a proxy for the deployment environment. The safety case then hinges on several claims: (1) the red team adequately elicits model capabilities to exfiltrate data, (2) control measures remain at least as effective in deployment, and (3) developers conservatively extrapolate model performance to predict the probability of data exfiltration in deployment. This safety case sketch is a step toward more concrete arguments that can be used to show that a dangerously capable LLM agent is safe to deploy.
title A sketch of an AI control safety case
topic Artificial Intelligence
Cryptography and Security
Software Engineering
url https://arxiv.org/abs/2501.17315