Holistic Safety and Responsibility Evaluations of Advanced AI Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Weidinger, Laura, Barnhart, Joslyn, Brennan, Jenny, Butterfield, Christina, Young, Susie, Hawkins, Will, Hendricks, Lisa Anne, Comanescu, Ramona, Chang, Oscar, Rodriguez, Mikel, Beroshi, Jennifer, Bloxwich, Dawn, Proleev, Lev, Chen, Jilin, Farquhar, Sebastian, Ho, Lewis, Gabriel, Iason, Dafoe, Allan, Isaac, William
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917646883094528
author Weidinger, Laura
Barnhart, Joslyn
Brennan, Jenny
Butterfield, Christina
Young, Susie
Hawkins, Will
Hendricks, Lisa Anne
Comanescu, Ramona
Chang, Oscar
Rodriguez, Mikel
Beroshi, Jennifer
Bloxwich, Dawn
Proleev, Lev
Chen, Jilin
Farquhar, Sebastian
Ho, Lewis
Gabriel, Iason
Dafoe, Allan
Isaac, William
author_facet Weidinger, Laura
Barnhart, Joslyn
Brennan, Jenny
Butterfield, Christina
Young, Susie
Hawkins, Will
Hendricks, Lisa Anne
Comanescu, Ramona
Chang, Oscar
Rodriguez, Mikel
Beroshi, Jennifer
Bloxwich, Dawn
Proleev, Lev
Chen, Jilin
Farquhar, Sebastian
Ho, Lewis
Gabriel, Iason
Dafoe, Allan
Isaac, William
contents Safety and responsibility evaluations of advanced AI models are a critical but developing field of research and practice. In the development of Google DeepMind's advanced AI models, we innovated on and applied a broad set of approaches to safety evaluation. In this report, we summarise and share elements of our evolving approach as well as lessons learned for a broad audience. Key lessons learned include: First, theoretical underpinnings and frameworks are invaluable to organise the breadth of risk domains, modalities, forms, metrics, and goals. Second, theory and practice of safety evaluation development each benefit from collaboration to clarify goals, methods and challenges, and facilitate the transfer of insights between different stakeholders and disciplines. Third, similar key methods, lessons, and institutions apply across the range of concerns in responsibility and safety - including established and emerging harms. For this reason it is important that a wide range of actors working on safety evaluation and safety research communities work together to develop, refine and implement novel evaluation approaches and best practices, rather than operating in silos. The report concludes with outlining the clear need to rapidly advance the science of evaluations, to integrate new evaluations into the development and governance of AI, to establish scientifically-grounded norms and standards, and to promote a robust evaluation ecosystem.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14068
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Holistic Safety and Responsibility Evaluations of Advanced AI Models
Weidinger, Laura
Barnhart, Joslyn
Brennan, Jenny
Butterfield, Christina
Young, Susie
Hawkins, Will
Hendricks, Lisa Anne
Comanescu, Ramona
Chang, Oscar
Rodriguez, Mikel
Beroshi, Jennifer
Bloxwich, Dawn
Proleev, Lev
Chen, Jilin
Farquhar, Sebastian
Ho, Lewis
Gabriel, Iason
Dafoe, Allan
Isaac, William
Artificial Intelligence
Machine Learning
Safety and responsibility evaluations of advanced AI models are a critical but developing field of research and practice. In the development of Google DeepMind's advanced AI models, we innovated on and applied a broad set of approaches to safety evaluation. In this report, we summarise and share elements of our evolving approach as well as lessons learned for a broad audience. Key lessons learned include: First, theoretical underpinnings and frameworks are invaluable to organise the breadth of risk domains, modalities, forms, metrics, and goals. Second, theory and practice of safety evaluation development each benefit from collaboration to clarify goals, methods and challenges, and facilitate the transfer of insights between different stakeholders and disciplines. Third, similar key methods, lessons, and institutions apply across the range of concerns in responsibility and safety - including established and emerging harms. For this reason it is important that a wide range of actors working on safety evaluation and safety research communities work together to develop, refine and implement novel evaluation approaches and best practices, rather than operating in silos. The report concludes with outlining the clear need to rapidly advance the science of evaluations, to integrate new evaluations into the development and governance of AI, to establish scientifically-grounded norms and standards, and to promote a robust evaluation ecosystem.
title Holistic Safety and Responsibility Evaluations of Advanced AI Models
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.14068