Adversarial Subspace Generation for Outlier Detection in High-Dimensional Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cribeiro-Ramallo, Jose, Matteucci, Federico, Enciu, Paul, Jenke, Alexander, Arzamasov, Vadim, Strufe, Thorsten, Böhm, Klemens
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908311304011776
author Cribeiro-Ramallo, Jose
Matteucci, Federico
Enciu, Paul
Jenke, Alexander
Arzamasov, Vadim
Strufe, Thorsten
Böhm, Klemens
author_facet Cribeiro-Ramallo, Jose
Matteucci, Federico
Enciu, Paul
Jenke, Alexander
Arzamasov, Vadim
Strufe, Thorsten
Böhm, Klemens
contents Outlier detection in high-dimensional tabular data is challenging since data is often distributed across multiple lower-dimensional subspaces -- a phenomenon known as the Multiple Views effect (MV). This effect led to a large body of research focused on mining such subspaces, known as subspace selection. However, as the precise nature of the MV effect was not well understood, traditional methods had to rely on heuristic-driven search schemes that struggle to accurately capture the true structure of the data. Properly identifying these subspaces is critical for unsupervised tasks such as outlier detection or clustering, where misrepresenting the underlying data structure can hinder the performance. We introduce Myopic Subspace Theory (MST), a new theoretical framework that mathematically formulates the Multiple Views effect and writes subspace selection as a stochastic optimization problem. Based on MST, we introduce V-GAN, a generative method trained to solve such an optimization problem. This approach avoids any exhaustive search over the feature space while ensuring that the intrinsic data structure is preserved. Experiments on 42 real-world datasets show that using V-GAN subspaces to build ensemble methods leads to a significant increase in one-class classification performance -- compared to existing subspace selection, feature selection, and embedding methods. Further experiments on synthetic data show that V-GAN identifies subspaces more accurately while scaling better than other relevant subspace selection methods. These results confirm the theoretical guarantees of our approach and also highlight its practical viability in high-dimensional settings.
format Preprint
id arxiv_https___arxiv_org_abs_2504_07522
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adversarial Subspace Generation for Outlier Detection in High-Dimensional Data
Cribeiro-Ramallo, Jose
Matteucci, Federico
Enciu, Paul
Jenke, Alexander
Arzamasov, Vadim
Strufe, Thorsten
Böhm, Klemens
Machine Learning
Artificial Intelligence
Statistics Theory
68T07
Outlier detection in high-dimensional tabular data is challenging since data is often distributed across multiple lower-dimensional subspaces -- a phenomenon known as the Multiple Views effect (MV). This effect led to a large body of research focused on mining such subspaces, known as subspace selection. However, as the precise nature of the MV effect was not well understood, traditional methods had to rely on heuristic-driven search schemes that struggle to accurately capture the true structure of the data. Properly identifying these subspaces is critical for unsupervised tasks such as outlier detection or clustering, where misrepresenting the underlying data structure can hinder the performance. We introduce Myopic Subspace Theory (MST), a new theoretical framework that mathematically formulates the Multiple Views effect and writes subspace selection as a stochastic optimization problem. Based on MST, we introduce V-GAN, a generative method trained to solve such an optimization problem. This approach avoids any exhaustive search over the feature space while ensuring that the intrinsic data structure is preserved. Experiments on 42 real-world datasets show that using V-GAN subspaces to build ensemble methods leads to a significant increase in one-class classification performance -- compared to existing subspace selection, feature selection, and embedding methods. Further experiments on synthetic data show that V-GAN identifies subspaces more accurately while scaling better than other relevant subspace selection methods. These results confirm the theoretical guarantees of our approach and also highlight its practical viability in high-dimensional settings.
title Adversarial Subspace Generation for Outlier Detection in High-Dimensional Data
topic Machine Learning
Artificial Intelligence
Statistics Theory
68T07
url https://arxiv.org/abs/2504.07522