Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mielniczuk, Jan, Wawrzeńczyk, Adam
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2312.02095
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908942923202560
author Mielniczuk, Jan
Wawrzeńczyk, Adam
author_facet Mielniczuk, Jan
Wawrzeńczyk, Adam
contents In the paper we argue that performance of the classifiers based on Empirical Risk Minimization (ERM) for positive unlabeled data, which are designed for case-control sampling scheme may significantly deteriorate when applied to a single-sample scenario. We reveal why their behavior depends, in all but very specific cases, on the scenario. Also, we introduce a single-sample case analogue of the popular non-negative risk classifier designed for case-control data and compare its performance with the original proposal. We show that the significant differences occur between them, especiall when half or more positive of observations are labeled. The opposite case when ERM minimizer designed for the case-control case is applied for single-sample data is also considered and similar conclusions are drawn. Taking into account difference of scenarios requires a sole, but crucial, change in the definition of the Empirical Risk.
format Preprint
id arxiv_https___arxiv_org_abs_2312_02095
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Single-sample versus case-control sampling scheme for Positive Unlabeled data: the story of two scenarios
Mielniczuk, Jan
Wawrzeńczyk, Adam
Machine Learning
In the paper we argue that performance of the classifiers based on Empirical Risk Minimization (ERM) for positive unlabeled data, which are designed for case-control sampling scheme may significantly deteriorate when applied to a single-sample scenario. We reveal why their behavior depends, in all but very specific cases, on the scenario. Also, we introduce a single-sample case analogue of the popular non-negative risk classifier designed for case-control data and compare its performance with the original proposal. We show that the significant differences occur between them, especiall when half or more positive of observations are labeled. The opposite case when ERM minimizer designed for the case-control case is applied for single-sample data is also considered and similar conclusions are drawn. Taking into account difference of scenarios requires a sole, but crucial, change in the definition of the Empirical Risk.
title Single-sample versus case-control sampling scheme for Positive Unlabeled data: the story of two scenarios
topic Machine Learning
url https://arxiv.org/abs/2312.02095