RegMixMatch: Optimizing Mixup Utilization in Semi-Supervised Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Han, Haorong, Yuan, Jidong, Wei, Chixuan, Yu, Zhongyang
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912332113772544
author Han, Haorong
Yuan, Jidong
Wei, Chixuan
Yu, Zhongyang
author_facet Han, Haorong
Yuan, Jidong
Wei, Chixuan
Yu, Zhongyang
contents Consistency regularization and pseudo-labeling have significantly advanced semi-supervised learning (SSL). Prior works have effectively employed Mixup for consistency regularization in SSL. However, our findings indicate that applying Mixup for consistency regularization may degrade SSL performance by compromising the purity of artificial labels. Moreover, most pseudo-labeling based methods utilize thresholding strategy to exclude low-confidence data, aiming to mitigate confirmation bias; however, this approach limits the utility of unlabeled samples. To address these challenges, we propose RegMixMatch, a novel framework that optimizes the use of Mixup with both high- and low-confidence samples in SSL. First, we introduce semi-supervised RegMixup, which effectively addresses reduced artificial labels purity by using both mixed samples and clean samples for training. Second, we develop a class-aware Mixup technique that integrates information from the top-2 predicted classes into low-confidence samples and their artificial labels, reducing the confirmation bias associated with these samples and enhancing their effective utilization. Experimental results demonstrate that RegMixMatch achieves state-of-the-art performance across various SSL benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10741
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RegMixMatch: Optimizing Mixup Utilization in Semi-Supervised Learning
Han, Haorong
Yuan, Jidong
Wei, Chixuan
Yu, Zhongyang
Machine Learning
Computer Vision and Pattern Recognition
Consistency regularization and pseudo-labeling have significantly advanced semi-supervised learning (SSL). Prior works have effectively employed Mixup for consistency regularization in SSL. However, our findings indicate that applying Mixup for consistency regularization may degrade SSL performance by compromising the purity of artificial labels. Moreover, most pseudo-labeling based methods utilize thresholding strategy to exclude low-confidence data, aiming to mitigate confirmation bias; however, this approach limits the utility of unlabeled samples. To address these challenges, we propose RegMixMatch, a novel framework that optimizes the use of Mixup with both high- and low-confidence samples in SSL. First, we introduce semi-supervised RegMixup, which effectively addresses reduced artificial labels purity by using both mixed samples and clean samples for training. Second, we develop a class-aware Mixup technique that integrates information from the top-2 predicted classes into low-confidence samples and their artificial labels, reducing the confirmation bias associated with these samples and enhancing their effective utilization. Experimental results demonstrate that RegMixMatch achieves state-of-the-art performance across various SSL benchmarks.
title RegMixMatch: Optimizing Mixup Utilization in Semi-Supervised Learning
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.10741