Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dorner, Florian E., Hardt, Moritz
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915922094063616
author Dorner, Florian E.
Hardt, Moritz
author_facet Dorner, Florian E.
Hardt, Moritz
contents We study how to best spend a budget of noisy labels to compare the accuracy of two binary classifiers. It's common practice to collect and aggregate multiple noisy labels for a given data point into a less noisy label via a majority vote. We prove a theorem that runs counter to conventional wisdom. If the goal is to identify the better of two classifiers, we show it's best to spend the budget on collecting a single label for more samples. Our result follows from a non-trivial application of Cramér's theorem, a staple in the theory of large deviations. We discuss the implications of our work for the design of machine learning benchmarks, where they overturn some time-honored recommendations. In addition, our results provide sample size bounds superior to what follows from Hoeffding's bound.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02249
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget
Dorner, Florian E.
Hardt, Moritz
Machine Learning
We study how to best spend a budget of noisy labels to compare the accuracy of two binary classifiers. It's common practice to collect and aggregate multiple noisy labels for a given data point into a less noisy label via a majority vote. We prove a theorem that runs counter to conventional wisdom. If the goal is to identify the better of two classifiers, we show it's best to spend the budget on collecting a single label for more samples. Our result follows from a non-trivial application of Cramér's theorem, a staple in the theory of large deviations. We discuss the implications of our work for the design of machine learning benchmarks, where they overturn some time-honored recommendations. In addition, our results provide sample size bounds superior to what follows from Hoeffding's bound.
title Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget
topic Machine Learning
url https://arxiv.org/abs/2402.02249