GPAI Evaluations Standards Taskforce: Towards Effective AI Governance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paskov, Patricia, Berglund, Lukas, Smith, Everett, Soder, Lisa
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913582482980864
author Paskov, Patricia
Berglund, Lukas
Smith, Everett
Soder, Lisa
author_facet Paskov, Patricia
Berglund, Lukas
Smith, Everett
Soder, Lisa
contents General-purpose AI evaluations have been proposed as a promising way of identifying and mitigating systemic risks posed by AI development and deployment. While GPAI evaluations play an increasingly central role in institutional decision- and policy-making -- including by way of the European Union AI Act's mandate to conduct evaluations on GPAI models presenting systemic risk -- no standards exist to date to promote their quality or legitimacy. To strengthen GPAI evaluations in the EU, which currently constitutes the first and only jurisdiction that mandates GPAI evaluations, we outline four desiderata for GPAI evaluations: internal validity, external validity, reproducibility, and portability. To uphold these desiderata in a dynamic environment of continuously evolving risks, we propose a dedicated EU GPAI Evaluation Standards Taskforce, to be housed within the bodies established by the EU AI Act. We outline the responsibilities of the Taskforce, specify the GPAI provider commitments that would facilitate Taskforce success, discuss the potential impact of the Taskforce on global AI governance, and address potential sources of failure that policymakers should heed.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13808
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GPAI Evaluations Standards Taskforce: Towards Effective AI Governance
Paskov, Patricia
Berglund, Lukas
Smith, Everett
Soder, Lisa
Computers and Society
General-purpose AI evaluations have been proposed as a promising way of identifying and mitigating systemic risks posed by AI development and deployment. While GPAI evaluations play an increasingly central role in institutional decision- and policy-making -- including by way of the European Union AI Act's mandate to conduct evaluations on GPAI models presenting systemic risk -- no standards exist to date to promote their quality or legitimacy. To strengthen GPAI evaluations in the EU, which currently constitutes the first and only jurisdiction that mandates GPAI evaluations, we outline four desiderata for GPAI evaluations: internal validity, external validity, reproducibility, and portability. To uphold these desiderata in a dynamic environment of continuously evolving risks, we propose a dedicated EU GPAI Evaluation Standards Taskforce, to be housed within the bodies established by the EU AI Act. We outline the responsibilities of the Taskforce, specify the GPAI provider commitments that would facilitate Taskforce success, discuss the potential impact of the Taskforce on global AI governance, and address potential sources of failure that policymakers should heed.
title GPAI Evaluations Standards Taskforce: Towards Effective AI Governance
topic Computers and Society
url https://arxiv.org/abs/2411.13808