Guardado en:
Detalles Bibliográficos
Autores principales: Mielniczuk, Jan, Rejchel, Wojciech, Teisseyre, Paweł
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2502.21194
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914585876889600
author Mielniczuk, Jan
Rejchel, Wojciech
Teisseyre, Paweł
author_facet Mielniczuk, Jan
Rejchel, Wojciech
Teisseyre, Paweł
contents We study estimation of a class prior for unlabeled target samples which possibly differs from that of source population. Moreover, it is assumed that the source data is partially observable: only samples from the positive class and from the whole population are available (PU learning scenario). We introduce a novel direct estimator of a class prior which avoids estimation of posterior probabilities in both populations and has a simple geometric interpretation. It is based on a distribution matching technique together with kernel embedding in a Reproducing Kernel Hilbert Space and is obtained as an explicit solution to an optimisation task. We establish its asymptotic consistency as well as an explicit non-asymptotic bound on its deviation from the unknown prior, which is calculable in practice. We study finite sample behaviour for synthetic and real data and show that the proposal works consistently on par or better than its competitors.
format Preprint
id arxiv_https___arxiv_org_abs_2502_21194
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prior shift estimation for positive unlabeled data through the lens of kernel embedding
Mielniczuk, Jan
Rejchel, Wojciech
Teisseyre, Paweł
Machine Learning
We study estimation of a class prior for unlabeled target samples which possibly differs from that of source population. Moreover, it is assumed that the source data is partially observable: only samples from the positive class and from the whole population are available (PU learning scenario). We introduce a novel direct estimator of a class prior which avoids estimation of posterior probabilities in both populations and has a simple geometric interpretation. It is based on a distribution matching technique together with kernel embedding in a Reproducing Kernel Hilbert Space and is obtained as an explicit solution to an optimisation task. We establish its asymptotic consistency as well as an explicit non-asymptotic bound on its deviation from the unknown prior, which is calculable in practice. We study finite sample behaviour for synthetic and real data and show that the proposal works consistently on par or better than its competitors.
title Prior shift estimation for positive unlabeled data through the lens of kernel embedding
topic Machine Learning
url https://arxiv.org/abs/2502.21194