On Benchmark Hacking in ML Contests: Modeling, Insights and Design

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Xiaoyun, Yu, Yang, Xu, Haifeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913058960441344
author Qiu, Xiaoyun
Yu, Yang
Xu, Haifeng
author_facet Qiu, Xiaoyun
Yu, Yang
Xu, Haifeng
contents Benchmark hacking refers to tuning a machine learning model to score highly on certain evaluation criteria without improving true generalization or faithfully solving the intended problem. We study this phenomenon in a generic machine learning contest, where each contestant chooses two types of effort: creative effort that improves model capability as desired by the contest host, and mechanistic effort that only improves the model's fitness to the particular task in contest without contributing to true generalization. We establish the existence of a symmetric monotone pure strategy equilibrium in this competition game. It also provides a natural definition of benchmark hacking in this strategic context by comparing a player's equilibrium effort allocation to that of a single-agent baseline scenario. Under our definition, contestants with types below certain threshold (low types) always engage in benchmark hacking, whereas those above the threshold do not. Furthermore, we show that more skewed reward structures (favoring top-ranked contestants) can elicit more desirable contest outcomes. We also provide empirical evidence to support our theoretical predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22230
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On Benchmark Hacking in ML Contests: Modeling, Insights and Design
Qiu, Xiaoyun
Yu, Yang
Xu, Haifeng
General Economics
Economics
Computer Science and Game Theory
Machine Learning
Benchmark hacking refers to tuning a machine learning model to score highly on certain evaluation criteria without improving true generalization or faithfully solving the intended problem. We study this phenomenon in a generic machine learning contest, where each contestant chooses two types of effort: creative effort that improves model capability as desired by the contest host, and mechanistic effort that only improves the model's fitness to the particular task in contest without contributing to true generalization. We establish the existence of a symmetric monotone pure strategy equilibrium in this competition game. It also provides a natural definition of benchmark hacking in this strategic context by comparing a player's equilibrium effort allocation to that of a single-agent baseline scenario. Under our definition, contestants with types below certain threshold (low types) always engage in benchmark hacking, whereas those above the threshold do not. Furthermore, we show that more skewed reward structures (favoring top-ranked contestants) can elicit more desirable contest outcomes. We also provide empirical evidence to support our theoretical predictions.
title On Benchmark Hacking in ML Contests: Modeling, Insights and Design
topic General Economics
Economics
Computer Science and Game Theory
Machine Learning
url https://arxiv.org/abs/2604.22230