Deprecating Benchmarks: Criteria and Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Joaquin, Ayrton San, Gipiškis, Rokas, Staufer, Leon, Gil, Ariel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908441631522816
author Joaquin, Ayrton San
Gipiškis, Rokas
Staufer, Leon
Gil, Ariel
author_facet Joaquin, Ayrton San
Gipiškis, Rokas
Staufer, Leon
Gil, Ariel
contents As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance on when and how benchmarks should be deprecated once they cease to effectively perform their purpose. This risks benchmark scores over-valuing model capabilities, or worse, obscuring capabilities and safety-washing. Based on a review of benchmarking practices, we propose criteria to decide when to fully or partially deprecate benchmarks, and a framework for deprecating benchmarks. Our work aims to advance the state of benchmarking towards rigorous and quality evaluations, especially for frontier models, and our recommendations are aimed to benefit benchmark developers, benchmark users, AI governance actors (across governments, academia, and industry panels), and policy makers.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06434
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deprecating Benchmarks: Criteria and Framework
Joaquin, Ayrton San
Gipiškis, Rokas
Staufer, Leon
Gil, Ariel
Computers and Society
Artificial Intelligence
Machine Learning
As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance on when and how benchmarks should be deprecated once they cease to effectively perform their purpose. This risks benchmark scores over-valuing model capabilities, or worse, obscuring capabilities and safety-washing. Based on a review of benchmarking practices, we propose criteria to decide when to fully or partially deprecate benchmarks, and a framework for deprecating benchmarks. Our work aims to advance the state of benchmarking towards rigorous and quality evaluations, especially for frontier models, and our recommendations are aimed to benefit benchmark developers, benchmark users, AI governance actors (across governments, academia, and industry panels), and policy makers.
title Deprecating Benchmarks: Criteria and Framework
topic Computers and Society
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.06434