Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Kevin L., Paskov, Patricia, Dev, Sunishchal, Byun, Michael J., Reuel, Anka, Roberts-Gaal, Xavier, Calcott, Rachel, Coxon, Evie, Deshpande, Chinmay
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912682293067776
author Wei, Kevin L.
Paskov, Patricia
Dev, Sunishchal
Byun, Michael J.
Reuel, Anka
Roberts-Gaal, Xavier
Calcott, Rachel
Coxon, Evie
Deshpande, Chinmay
author_facet Wei, Kevin L.
Paskov, Patricia
Dev, Sunishchal
Byun, Michael J.
Reuel, Anka
Roberts-Gaal, Xavier
Calcott, Rachel
Coxon, Evie
Deshpande, Chinmay
contents In this position paper, we argue that human baselines in foundation model evaluations must be more rigorous and more transparent to enable meaningful comparisons of human vs. AI performance, and we provide recommendations and a reporting checklist towards this end. Human performance baselines are vital for the machine learning community, downstream users, and policymakers to interpret AI evaluations. Models are often claimed to achieve "super-human" performance, but existing baselining methods are neither sufficiently rigorous nor sufficiently well-documented to robustly measure and assess performance differences. Based on a meta-review of the measurement theory and AI evaluation literatures, we derive a framework with recommendations for designing, executing, and reporting human baselines. We synthesize our recommendations into a checklist that we use to systematically review 115 human baselines (studies) in foundation model evaluations and thus identify shortcomings in existing baselining methods; our checklist can also assist researchers in conducting human baselines and reporting results. We hope our work can advance more rigorous AI evaluation practices that can better serve both the research community and policymakers. Data is available at: https://github.com/kevinlwei/human-baselines
format Preprint
id arxiv_https___arxiv_org_abs_2506_13776
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
Wei, Kevin L.
Paskov, Patricia
Dev, Sunishchal
Byun, Michael J.
Reuel, Anka
Roberts-Gaal, Xavier
Calcott, Rachel
Coxon, Evie
Deshpande, Chinmay
Artificial Intelligence
Computers and Society
Human-Computer Interaction
In this position paper, we argue that human baselines in foundation model evaluations must be more rigorous and more transparent to enable meaningful comparisons of human vs. AI performance, and we provide recommendations and a reporting checklist towards this end. Human performance baselines are vital for the machine learning community, downstream users, and policymakers to interpret AI evaluations. Models are often claimed to achieve "super-human" performance, but existing baselining methods are neither sufficiently rigorous nor sufficiently well-documented to robustly measure and assess performance differences. Based on a meta-review of the measurement theory and AI evaluation literatures, we derive a framework with recommendations for designing, executing, and reporting human baselines. We synthesize our recommendations into a checklist that we use to systematically review 115 human baselines (studies) in foundation model evaluations and thus identify shortcomings in existing baselining methods; our checklist can also assist researchers in conducting human baselines and reporting results. We hope our work can advance more rigorous AI evaluation practices that can better serve both the research community and policymakers. Data is available at: https://github.com/kevinlwei/human-baselines
title Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
topic Artificial Intelligence
Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2506.13776