Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Zhimin, Wang, Zehao, Bangash, Abdul Ali, Adams, Bram, Hassan, Ahmed E.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917526616670208
author Zhao, Zhimin
Wang, Zehao
Bangash, Abdul Ali
Adams, Bram
Hassan, Ahmed E.
author_facet Zhao, Zhimin
Wang, Zehao
Bangash, Abdul Ali
Adams, Bram
Hassan, Ahmed E.
contents Evaluation harnesses are software systems that orchestrate model evaluation by managing model invocation, data loading, metric computation, and result reporting. Despite their critical role in machine learning infrastructure, their operational challenges and engineering concerns have received limited attention so far. We present an empirical study of 57 evaluation harnesses, deriving a five-stage harness model and classifying 16,560 issues by workflow stage and root cause. Most harness operational challenges concentrate in the Specification stage (41.4% of issues), where harnesses integrate external models, datasets, and scoring judges. The three most frequent root causes of operational challenges are unimplemented features (24.3%), documentation gaps (20.3%), and missing input validation (17.2%), which together account for 61.7% of classified issues, spanning both defects in existing functionality and capability gaps that block intended workflows. Root causes also vary by workflow stage: environment incompatibility and external dependency breakage account for 36.2% of provisioning issues, whereas algorithmic error (25.9%) and validation gap (22.5%) dominate assessment issues. Together, these contributions establish an empirical foundation for treating evaluation engineering as a distinct software engineering concern.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24213
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
Zhao, Zhimin
Wang, Zehao
Bangash, Abdul Ali
Adams, Bram
Hassan, Ahmed E.
Software Engineering
Artificial Intelligence
Machine Learning
Evaluation harnesses are software systems that orchestrate model evaluation by managing model invocation, data loading, metric computation, and result reporting. Despite their critical role in machine learning infrastructure, their operational challenges and engineering concerns have received limited attention so far. We present an empirical study of 57 evaluation harnesses, deriving a five-stage harness model and classifying 16,560 issues by workflow stage and root cause. Most harness operational challenges concentrate in the Specification stage (41.4% of issues), where harnesses integrate external models, datasets, and scoring judges. The three most frequent root causes of operational challenges are unimplemented features (24.3%), documentation gaps (20.3%), and missing input validation (17.2%), which together account for 61.7% of classified issues, spanning both defects in existing functionality and capability gaps that block intended workflows. Root causes also vary by workflow stage: environment incompatibility and external dependency breakage account for 36.2% of provisioning issues, whereas algorithmic error (25.9%) and validation gap (22.5%) dominate assessment issues. Together, these contributions establish an empirical foundation for treating evaluation engineering as a distinct software engineering concern.
title Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
topic Software Engineering
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.24213