Learning from Uncertain Similarity and Unlabeled Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Meng, Li, Zhongnian, Ying, Peng, Xu, Xinzheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909788817850368
author Wei, Meng
Li, Zhongnian
Ying, Peng
Xu, Xinzheng
author_facet Wei, Meng
Li, Zhongnian
Ying, Peng
Xu, Xinzheng
contents Existing similarity-based weakly supervised learning approaches often rely on precise similarity annotations between data pairs, which may inadvertently expose sensitive label information and raise privacy risks. To mitigate this issue, we propose Uncertain Similarity and Unlabeled Learning (USimUL), a novel framework where each similarity pair is embedded with an uncertainty component to reduce label leakage. In this paper, we propose an unbiased risk estimator that learns from uncertain similarity and unlabeled data. Additionally, we theoretically prove that the estimator achieves statistically optimal parametric convergence rates. Extensive experiments on both benchmark and real-world datasets show that our method achieves superior classification performance compared to conventional similarity-based approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11984
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning from Uncertain Similarity and Unlabeled Data
Wei, Meng
Li, Zhongnian
Ying, Peng
Xu, Xinzheng
Machine Learning
Existing similarity-based weakly supervised learning approaches often rely on precise similarity annotations between data pairs, which may inadvertently expose sensitive label information and raise privacy risks. To mitigate this issue, we propose Uncertain Similarity and Unlabeled Learning (USimUL), a novel framework where each similarity pair is embedded with an uncertainty component to reduce label leakage. In this paper, we propose an unbiased risk estimator that learns from uncertain similarity and unlabeled data. Additionally, we theoretically prove that the estimator achieves statistically optimal parametric convergence rates. Extensive experiments on both benchmark and real-world datasets show that our method achieves superior classification performance compared to conventional similarity-based approaches.
title Learning from Uncertain Similarity and Unlabeled Data
topic Machine Learning
url https://arxiv.org/abs/2509.11984