Preliminary suggestions for rigorous GPAI model evaluations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paskov, Patricia, Byun, Michael J., Wei, Kevin, Webster, Toby
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908494196637696
author Paskov, Patricia
Byun, Michael J.
Wei, Kevin
Webster, Toby
author_facet Paskov, Patricia
Byun, Michael J.
Wei, Kevin
Webster, Toby
contents This document presents a preliminary compilation of general-purpose AI (GPAI) evaluation practices that may promote internal validity, external validity and reproducibility. It includes suggestions for human uplift studies and benchmark evaluations, as well as cross-cutting suggestions that may apply to many different evaluation types. Suggestions are organised across four stages in the evaluation life cycle: design, implementation, execution and documentation. Drawing from established practices in machine learning, statistics, psychology, economics, biology and other fields recognised to have important lessons for AI evaluation, these suggestions seek to contribute to the conversation on the nascent and evolving field of the science of GPAI evaluations. The intended audience of this document includes providers of GPAI models presenting systemic risk (GPAISR), for whom the EU AI Act lays out specific evaluation requirements; third-party evaluators; policymakers assessing the rigour of evaluations; and academic researchers developing or conducting GPAI evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00875
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Preliminary suggestions for rigorous GPAI model evaluations
Paskov, Patricia
Byun, Michael J.
Wei, Kevin
Webster, Toby
Computers and Society
Artificial Intelligence
This document presents a preliminary compilation of general-purpose AI (GPAI) evaluation practices that may promote internal validity, external validity and reproducibility. It includes suggestions for human uplift studies and benchmark evaluations, as well as cross-cutting suggestions that may apply to many different evaluation types. Suggestions are organised across four stages in the evaluation life cycle: design, implementation, execution and documentation. Drawing from established practices in machine learning, statistics, psychology, economics, biology and other fields recognised to have important lessons for AI evaluation, these suggestions seek to contribute to the conversation on the nascent and evolving field of the science of GPAI evaluations. The intended audience of this document includes providers of GPAI models presenting systemic risk (GPAISR), for whom the EU AI Act lays out specific evaluation requirements; third-party evaluators; policymakers assessing the rigour of evaluations; and academic researchers developing or conducting GPAI evaluations.
title Preliminary suggestions for rigorous GPAI model evaluations
topic Computers and Society
Artificial Intelligence
url https://arxiv.org/abs/2508.00875