Learning from Ambiguous Data with Hard Labels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Zeke, He, Zheng, Lu, Nan, Bai, Lichen, Li, Bao, Yang, Shuo, Sun, Mingming, Li, Ping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915095034986496
author Xie, Zeke
He, Zheng
Lu, Nan
Bai, Lichen
Li, Bao
Yang, Shuo
Sun, Mingming
Li, Ping
author_facet Xie, Zeke
He, Zheng
Lu, Nan
Bai, Lichen
Li, Bao
Yang, Shuo
Sun, Mingming
Li, Ping
contents Real-world data often contains intrinsic ambiguity that the common single-hard-label annotation paradigm ignores. Standard training using ambiguous data with these hard labels may produce overly confident models and thus leading to poor generalization. In this paper, we propose a novel framework called Quantized Label Learning (QLL) to alleviate this issue. First, we formulate QLL as learning from (very) ambiguous data with hard labels: ideally, each ambiguous instance should be associated with a ground-truth soft-label distribution describing its corresponding probabilistic weight in each class, however, this is usually not accessible; in practice, we can only observe a quantized label, i.e., a hard label sampled (quantized) from the corresponding ground-truth soft-label distribution, of each instance, which can be seen as a biased approximation of the ground-truth soft-label. Second, we propose a Class-wise Positive-Unlabeled (CPU) risk estimator that allows us to train accurate classifiers from only ambiguous data with quantized labels. Third, to simulate ambiguous datasets with quantized labels in the real world, we design a mixing-based ambiguous data generation procedure for empirical evaluation. Experiments demonstrate that our CPU method can significantly improve model generalization performance and outperform the baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01844
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning from Ambiguous Data with Hard Labels
Xie, Zeke
He, Zheng
Lu, Nan
Bai, Lichen
Li, Bao
Yang, Shuo
Sun, Mingming
Li, Ping
Machine Learning
Real-world data often contains intrinsic ambiguity that the common single-hard-label annotation paradigm ignores. Standard training using ambiguous data with these hard labels may produce overly confident models and thus leading to poor generalization. In this paper, we propose a novel framework called Quantized Label Learning (QLL) to alleviate this issue. First, we formulate QLL as learning from (very) ambiguous data with hard labels: ideally, each ambiguous instance should be associated with a ground-truth soft-label distribution describing its corresponding probabilistic weight in each class, however, this is usually not accessible; in practice, we can only observe a quantized label, i.e., a hard label sampled (quantized) from the corresponding ground-truth soft-label distribution, of each instance, which can be seen as a biased approximation of the ground-truth soft-label. Second, we propose a Class-wise Positive-Unlabeled (CPU) risk estimator that allows us to train accurate classifiers from only ambiguous data with quantized labels. Third, to simulate ambiguous datasets with quantized labels in the real world, we design a mixing-based ambiguous data generation procedure for empirical evaluation. Experiments demonstrate that our CPU method can significantly improve model generalization performance and outperform the baselines.
title Learning from Ambiguous Data with Hard Labels
topic Machine Learning
url https://arxiv.org/abs/2501.01844