Robust Semi-Supervised Learning in Open Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Lan-Zhe, Jia, Lin-Han, Shao, Jie-Jing, Li, Yu-Feng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917878691790848
author Guo, Lan-Zhe
Jia, Lin-Han
Shao, Jie-Jing
Li, Yu-Feng
author_facet Guo, Lan-Zhe
Jia, Lin-Han
Shao, Jie-Jing
Li, Yu-Feng
contents Semi-supervised learning (SSL) aims to improve performance by exploiting unlabeled data when labels are scarce. Conventional SSL studies typically assume close environments where important factors (e.g., label, feature, distribution) between labeled and unlabeled data are consistent. However, more practical tasks involve open environments where important factors between labeled and unlabeled data are inconsistent. It has been reported that exploiting inconsistent unlabeled data causes severe performance degradation, even worse than the simple supervised learning baseline. Manually verifying the quality of unlabeled data is not desirable, therefore, it is important to study robust SSL with inconsistent unlabeled data in open environments. This paper briefly introduces some advances in this line of research, focusing on techniques concerning label, feature, and data distribution inconsistency in SSL, and presents the evaluation benchmarks. Open research problems are also discussed for reference purposes.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18256
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robust Semi-Supervised Learning in Open Environments
Guo, Lan-Zhe
Jia, Lin-Han
Shao, Jie-Jing
Li, Yu-Feng
Machine Learning
Artificial Intelligence
Semi-supervised learning (SSL) aims to improve performance by exploiting unlabeled data when labels are scarce. Conventional SSL studies typically assume close environments where important factors (e.g., label, feature, distribution) between labeled and unlabeled data are consistent. However, more practical tasks involve open environments where important factors between labeled and unlabeled data are inconsistent. It has been reported that exploiting inconsistent unlabeled data causes severe performance degradation, even worse than the simple supervised learning baseline. Manually verifying the quality of unlabeled data is not desirable, therefore, it is important to study robust SSL with inconsistent unlabeled data in open environments. This paper briefly introduces some advances in this line of research, focusing on techniques concerning label, feature, and data distribution inconsistency in SSL, and presents the evaluation benchmarks. Open research problems are also discussed for reference purposes.
title Robust Semi-Supervised Learning in Open Environments
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2412.18256