Black-Box Access is Insufficient for Rigorous AI Audits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Casper, Stephen, Ezell, Carson, Siegmann, Charlotte, Kolt, Noam, Curtis, Taylor Lynn, Bucknall, Benjamin, Haupt, Andreas, Wei, Kevin, Scheurer, Jérémy, Hobbhahn, Marius, Sharkey, Lee, Krishna, Satyapriya, Von Hagen, Marvin, Alberti, Silas, Chan, Alan, Sun, Qinyi, Gerovitch, Michael, Bau, David, Tegmark, Max, Krueger, David, Hadfield-Menell, Dylan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911910352388096
author Casper, Stephen
Ezell, Carson
Siegmann, Charlotte
Kolt, Noam
Curtis, Taylor Lynn
Bucknall, Benjamin
Haupt, Andreas
Wei, Kevin
Scheurer, Jérémy
Hobbhahn, Marius
Sharkey, Lee
Krishna, Satyapriya
Von Hagen, Marvin
Alberti, Silas
Chan, Alan
Sun, Qinyi
Gerovitch, Michael
Bau, David
Tegmark, Max
Krueger, David
Hadfield-Menell, Dylan
author_facet Casper, Stephen
Ezell, Carson
Siegmann, Charlotte
Kolt, Noam
Curtis, Taylor Lynn
Bucknall, Benjamin
Haupt, Andreas
Wei, Kevin
Scheurer, Jérémy
Hobbhahn, Marius
Sharkey, Lee
Krishna, Satyapriya
Von Hagen, Marvin
Alberti, Silas
Chan, Alan
Sun, Qinyi
Gerovitch, Michael
Bau, David
Tegmark, Max
Krueger, David
Hadfield-Menell, Dylan
contents External audits of AI systems are increasingly recognized as a key mechanism for AI governance. The effectiveness of an audit, however, depends on the degree of access granted to auditors. Recent audits of state-of-the-art AI systems have primarily relied on black-box access, in which auditors can only query the system and observe its outputs. However, white-box access to the system's inner workings (e.g., weights, activations, gradients) allows an auditor to perform stronger attacks, more thoroughly interpret models, and conduct fine-tuning. Meanwhile, outside-the-box access to training and deployment information (e.g., methodology, code, documentation, data, deployment details, findings from internal evaluations) allows auditors to scrutinize the development process and design more targeted evaluations. In this paper, we examine the limitations of black-box audits and the advantages of white- and outside-the-box audits. We also discuss technical, physical, and legal safeguards for performing these audits with minimal security risks. Given that different forms of access can lead to very different levels of evaluation, we conclude that (1) transparency regarding the access and methods used by auditors is necessary to properly interpret audit results, and (2) white- and outside-the-box access allow for substantially more scrutiny than black-box access alone.
format Preprint
id arxiv_https___arxiv_org_abs_2401_14446
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Black-Box Access is Insufficient for Rigorous AI Audits
Casper, Stephen
Ezell, Carson
Siegmann, Charlotte
Kolt, Noam
Curtis, Taylor Lynn
Bucknall, Benjamin
Haupt, Andreas
Wei, Kevin
Scheurer, Jérémy
Hobbhahn, Marius
Sharkey, Lee
Krishna, Satyapriya
Von Hagen, Marvin
Alberti, Silas
Chan, Alan
Sun, Qinyi
Gerovitch, Michael
Bau, David
Tegmark, Max
Krueger, David
Hadfield-Menell, Dylan
Computers and Society
Artificial Intelligence
Cryptography and Security
External audits of AI systems are increasingly recognized as a key mechanism for AI governance. The effectiveness of an audit, however, depends on the degree of access granted to auditors. Recent audits of state-of-the-art AI systems have primarily relied on black-box access, in which auditors can only query the system and observe its outputs. However, white-box access to the system's inner workings (e.g., weights, activations, gradients) allows an auditor to perform stronger attacks, more thoroughly interpret models, and conduct fine-tuning. Meanwhile, outside-the-box access to training and deployment information (e.g., methodology, code, documentation, data, deployment details, findings from internal evaluations) allows auditors to scrutinize the development process and design more targeted evaluations. In this paper, we examine the limitations of black-box audits and the advantages of white- and outside-the-box audits. We also discuss technical, physical, and legal safeguards for performing these audits with minimal security risks. Given that different forms of access can lead to very different levels of evaluation, we conclude that (1) transparency regarding the access and methods used by auditors is necessary to properly interpret audit results, and (2) white- and outside-the-box access allow for substantially more scrutiny than black-box access alone.
title Black-Box Access is Insufficient for Rigorous AI Audits
topic Computers and Society
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2401.14446