Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915922094063616 |
|---|---|
| author | Dorner, Florian E. Hardt, Moritz |
| author_facet | Dorner, Florian E. Hardt, Moritz |
| contents | We study how to best spend a budget of noisy labels to compare the accuracy of two binary classifiers. It's common practice to collect and aggregate multiple noisy labels for a given data point into a less noisy label via a majority vote. We prove a theorem that runs counter to conventional wisdom. If the goal is to identify the better of two classifiers, we show it's best to spend the budget on collecting a single label for more samples. Our result follows from a non-trivial application of Cramér's theorem, a staple in the theory of large deviations. We discuss the implications of our work for the design of machine learning benchmarks, where they overturn some time-honored recommendations. In addition, our results provide sample size bounds superior to what follows from Hoeffding's bound. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_02249 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget Dorner, Florian E. Hardt, Moritz Machine Learning We study how to best spend a budget of noisy labels to compare the accuracy of two binary classifiers. It's common practice to collect and aggregate multiple noisy labels for a given data point into a less noisy label via a majority vote. We prove a theorem that runs counter to conventional wisdom. If the goal is to identify the better of two classifiers, we show it's best to spend the budget on collecting a single label for more samples. Our result follows from a non-trivial application of Cramér's theorem, a staple in the theory of large deviations. We discuss the implications of our work for the design of machine learning benchmarks, where they overturn some time-honored recommendations. In addition, our results provide sample size bounds superior to what follows from Hoeffding's bound. |
| title | Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2402.02249 |