Deep Positive-Unlabeled Anomaly Detection for Contaminated Unlabeled Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Takahashi, Hiroshi, Iwata, Tomoharu, Kumagai, Atsutoshi, Yamanaka, Yuuki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909483391778816
author Takahashi, Hiroshi
Iwata, Tomoharu
Kumagai, Atsutoshi
Yamanaka, Yuuki
author_facet Takahashi, Hiroshi
Iwata, Tomoharu
Kumagai, Atsutoshi
Yamanaka, Yuuki
contents Semi-supervised anomaly detection, which aims to improve the anomaly detection performance by using a small amount of labeled anomaly data in addition to unlabeled data, has attracted attention. Existing semi-supervised approaches assume that most unlabeled data are normal, and train anomaly detectors by minimizing the anomaly scores for the unlabeled data while maximizing those for the labeled anomaly data. However, in practice, the unlabeled data are often contaminated with anomalies. This weakens the effect of maximizing the anomaly scores for anomalies, and prevents us from improving the detection performance. To solve this problem, we propose the deep positive-unlabeled anomaly detection framework, which integrates positive-unlabeled learning with deep anomaly detection models such as autoencoders and deep support vector data descriptions. Our approach enables the approximation of anomaly scores for normal data using the unlabeled data and the labeled anomaly data. Therefore, without labeled normal data, our approach can train anomaly detectors by minimizing the anomaly scores for normal data while maximizing those for the labeled anomaly data. Experiments on various datasets show that our approach achieves better detection performance than existing approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18929
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Deep Positive-Unlabeled Anomaly Detection for Contaminated Unlabeled Data
Takahashi, Hiroshi
Iwata, Tomoharu
Kumagai, Atsutoshi
Yamanaka, Yuuki
Machine Learning
Artificial Intelligence
Semi-supervised anomaly detection, which aims to improve the anomaly detection performance by using a small amount of labeled anomaly data in addition to unlabeled data, has attracted attention. Existing semi-supervised approaches assume that most unlabeled data are normal, and train anomaly detectors by minimizing the anomaly scores for the unlabeled data while maximizing those for the labeled anomaly data. However, in practice, the unlabeled data are often contaminated with anomalies. This weakens the effect of maximizing the anomaly scores for anomalies, and prevents us from improving the detection performance. To solve this problem, we propose the deep positive-unlabeled anomaly detection framework, which integrates positive-unlabeled learning with deep anomaly detection models such as autoencoders and deep support vector data descriptions. Our approach enables the approximation of anomaly scores for normal data using the unlabeled data and the labeled anomaly data. Therefore, without labeled normal data, our approach can train anomaly detectors by minimizing the anomaly scores for normal data while maximizing those for the labeled anomaly data. Experiments on various datasets show that our approach achieves better detection performance than existing approaches.
title Deep Positive-Unlabeled Anomaly Detection for Contaminated Unlabeled Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.18929