Improving the Validity and Practical Usefulness of AI/ML Evaluations Using an Estimands Framework

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Binette, Olivier, Reiter, Jerome P.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929387183538176
author Binette, Olivier
Reiter, Jerome P.
author_facet Binette, Olivier
Reiter, Jerome P.
contents Commonly, AI or machine learning (ML) models are evaluated on benchmark datasets. This practice supports innovative methodological research, but benchmark performance can be poorly correlated with performance in real-world applications -- a construct validity issue. To improve the validity and practical usefulness of evaluations, we propose using an estimands framework adapted from international clinical trials guidelines. This framework provides a systematic structure for inference and reporting in evaluations, emphasizing the importance of a well-defined estimation target. We illustrate our proposal on examples of commonly used evaluation methodologies - involving cross-validation, clustering evaluation, and LLM benchmarking - that can lead to incorrect rankings of competing models (rank reversals) with high probability, even when performance differences are large. We demonstrate how the estimands framework can help uncover underlying issues, their causes, and potential solutions. Ultimately, we believe this framework can improve the validity of evaluations through better-aligned inference, and help decision-makers and model users interpret reported results more effectively.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10366
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving the Validity and Practical Usefulness of AI/ML Evaluations Using an Estimands Framework
Binette, Olivier
Reiter, Jerome P.
Machine Learning
Applications
Methodology
Commonly, AI or machine learning (ML) models are evaluated on benchmark datasets. This practice supports innovative methodological research, but benchmark performance can be poorly correlated with performance in real-world applications -- a construct validity issue. To improve the validity and practical usefulness of evaluations, we propose using an estimands framework adapted from international clinical trials guidelines. This framework provides a systematic structure for inference and reporting in evaluations, emphasizing the importance of a well-defined estimation target. We illustrate our proposal on examples of commonly used evaluation methodologies - involving cross-validation, clustering evaluation, and LLM benchmarking - that can lead to incorrect rankings of competing models (rank reversals) with high probability, even when performance differences are large. We demonstrate how the estimands framework can help uncover underlying issues, their causes, and potential solutions. Ultimately, we believe this framework can improve the validity of evaluations through better-aligned inference, and help decision-makers and model users interpret reported results more effectively.
title Improving the Validity and Practical Usefulness of AI/ML Evaluations Using an Estimands Framework
topic Machine Learning
Applications
Methodology
url https://arxiv.org/abs/2406.10366