Anomaly Detection with Adaptive and Aggressive Rejection for Contaminated Training Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Jungi, Kim, Jungkwon, Zhang, Chi, Yoo, Kwangsun, Byun, Seok-Joo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914171871821824
author Lee, Jungi
Kim, Jungkwon
Zhang, Chi
Yoo, Kwangsun
Byun, Seok-Joo
author_facet Lee, Jungi
Kim, Jungkwon
Zhang, Chi
Yoo, Kwangsun
Byun, Seok-Joo
contents Handling contaminated data poses a critical challenge in anomaly detection, as traditional models assume training on purely normal data. Conventional methods mitigate contamination by relying on fixed contamination ratios, but discrepancies between assumed and actual ratios can severely degrade performance, especially in noisy environments where normal and abnormal data distributions overlap. To address these limitations, we propose Adaptive and Aggressive Rejection (AAR), a novel method that dynamically excludes anomalies using a modified z-score and Gaussian mixture model-based thresholds. AAR effectively balances the trade-off between preserving normal data and excluding anomalies by integrating hard and soft rejection strategies. Extensive experiments on two image datasets and thirty tabular datasets demonstrate that AAR outperforms the state-of-the-art method by 0.041 AUROC. By providing a scalable and reliable solution, AAR enhances robustness against contaminated datasets, paving the way for broader real-world applications in domains such as security and healthcare.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21378
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Anomaly Detection with Adaptive and Aggressive Rejection for Contaminated Training Data
Lee, Jungi
Kim, Jungkwon
Zhang, Chi
Yoo, Kwangsun
Byun, Seok-Joo
Machine Learning
Artificial Intelligence
Handling contaminated data poses a critical challenge in anomaly detection, as traditional models assume training on purely normal data. Conventional methods mitigate contamination by relying on fixed contamination ratios, but discrepancies between assumed and actual ratios can severely degrade performance, especially in noisy environments where normal and abnormal data distributions overlap. To address these limitations, we propose Adaptive and Aggressive Rejection (AAR), a novel method that dynamically excludes anomalies using a modified z-score and Gaussian mixture model-based thresholds. AAR effectively balances the trade-off between preserving normal data and excluding anomalies by integrating hard and soft rejection strategies. Extensive experiments on two image datasets and thirty tabular datasets demonstrate that AAR outperforms the state-of-the-art method by 0.041 AUROC. By providing a scalable and reliable solution, AAR enhances robustness against contaminated datasets, paving the way for broader real-world applications in domains such as security and healthcare.
title Anomaly Detection with Adaptive and Aggressive Rejection for Contaminated Training Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.21378