Affirmative safety: An approach to risk management for high-risk AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wasil, Akash R., Clymer, Joshua, Krueger, David, Dardaman, Emily, Campos, Simeon, Murphy, Evan R.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913400897929216
author Wasil, Akash R.
Clymer, Joshua
Krueger, David
Dardaman, Emily
Campos, Simeon
Murphy, Evan R.
author_facet Wasil, Akash R.
Clymer, Joshua
Krueger, David
Dardaman, Emily
Campos, Simeon
Murphy, Evan R.
contents Prominent AI experts have suggested that companies developing high-risk AI systems should be required to show that such systems are safe before they can be developed or deployed. The goal of this paper is to expand on this idea and explore its implications for risk management. We argue that entities developing or deploying high-risk AI systems should be required to present evidence of affirmative safety: a proactive case that their activities keep risks below acceptable thresholds. We begin the paper by highlighting global security risks from AI that have been acknowledged by AI experts and world governments. Next, we briefly describe principles of risk management from other high-risk fields (e.g., nuclear safety). Then, we propose a risk management approach for advanced AI in which model developers must provide evidence that their activities keep certain risks below regulator-set thresholds. As a first step toward understanding what affirmative safety cases should include, we illustrate how certain kinds of technical evidence and operational evidence can support an affirmative safety case. In the technical section, we discuss behavioral evidence (evidence about model outputs), cognitive evidence (evidence about model internals), and developmental evidence (evidence about the training process). In the operational section, we offer examples of organizational practices that could contribute to affirmative safety cases: information security practices, safety culture, and emergency response capacity. Finally, we briefly compare our approach to the NIST AI Risk Management Framework. Overall, we hope our work contributes to ongoing discussions about national and global security risks posed by AI and regulatory approaches to address these risks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15371
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Affirmative safety: An approach to risk management for high-risk AI
Wasil, Akash R.
Clymer, Joshua
Krueger, David
Dardaman, Emily
Campos, Simeon
Murphy, Evan R.
Computers and Society
Artificial Intelligence
Prominent AI experts have suggested that companies developing high-risk AI systems should be required to show that such systems are safe before they can be developed or deployed. The goal of this paper is to expand on this idea and explore its implications for risk management. We argue that entities developing or deploying high-risk AI systems should be required to present evidence of affirmative safety: a proactive case that their activities keep risks below acceptable thresholds. We begin the paper by highlighting global security risks from AI that have been acknowledged by AI experts and world governments. Next, we briefly describe principles of risk management from other high-risk fields (e.g., nuclear safety). Then, we propose a risk management approach for advanced AI in which model developers must provide evidence that their activities keep certain risks below regulator-set thresholds. As a first step toward understanding what affirmative safety cases should include, we illustrate how certain kinds of technical evidence and operational evidence can support an affirmative safety case. In the technical section, we discuss behavioral evidence (evidence about model outputs), cognitive evidence (evidence about model internals), and developmental evidence (evidence about the training process). In the operational section, we offer examples of organizational practices that could contribute to affirmative safety cases: information security practices, safety culture, and emergency response capacity. Finally, we briefly compare our approach to the NIST AI Risk Management Framework. Overall, we hope our work contributes to ongoing discussions about national and global security risks posed by AI and regulatory approaches to address these risks.
title Affirmative safety: An approach to risk management for high-risk AI
topic Computers and Society
Artificial Intelligence
url https://arxiv.org/abs/2406.15371