Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Griffin, Charlie, Thomson, Louis, Shlegeris, Buck, Abate, Alessandro
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913097603612672
author Griffin, Charlie
Thomson, Louis
Shlegeris, Buck
Abate, Alessandro
author_facet Griffin, Charlie
Thomson, Louis
Shlegeris, Buck
Abate, Alessandro
contents To evaluate the safety and usefulness of deployment protocols for untrusted AIs, AI Control uses a red-teaming exercise played between a protocol designer and an adversary. This paper introduces AI-Control Games, a formal decision-making model of the red-teaming exercise as a multi-objective, partially observable, stochastic game. We also introduce reductions from AI-Control Games to a special case of zero-sum partially observable stochastic games that allow us to leverage existing algorithms to find Pareto-optimal protocols. We apply our formalism to model, evaluate and synthesise protocols for deploying untrusted language models as programming assistants, focusing on Trusted Monitoring protocols, which use weaker language models and limited human assistance. To demonstrate the utility of our formalism, we show improvements over empirical studies in existing settings, evaluate protocols in new settings, and analyse how modelling assumptions affect the safety and usefulness of protocols. Finally, we leverage our formalism to precisely describe some of the implicit assumptions in prior control work.
format Preprint
id arxiv_https___arxiv_org_abs_2409_07985
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
Griffin, Charlie
Thomson, Louis
Shlegeris, Buck
Abate, Alessandro
Artificial Intelligence
Machine Learning
To evaluate the safety and usefulness of deployment protocols for untrusted AIs, AI Control uses a red-teaming exercise played between a protocol designer and an adversary. This paper introduces AI-Control Games, a formal decision-making model of the red-teaming exercise as a multi-objective, partially observable, stochastic game. We also introduce reductions from AI-Control Games to a special case of zero-sum partially observable stochastic games that allow us to leverage existing algorithms to find Pareto-optimal protocols. We apply our formalism to model, evaluate and synthesise protocols for deploying untrusted language models as programming assistants, focusing on Trusted Monitoring protocols, which use weaker language models and limited human assistance. To demonstrate the utility of our formalism, we show improvements over empirical studies in existing settings, evaluate protocols in new settings, and analyse how modelling assumptions affect the safety and usefulness of protocols. Finally, we leverage our formalism to precisely describe some of the implicit assumptions in prior control work.
title Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2409.07985