Dataset-Adaptive Dimensionality Reduction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jeon, Hyeon, Park, Jeongin, Lee, Soohyun, Kim, Dae Hyun, Shin, Sungbok, Seo, Jinwook
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912816154279936
author Jeon, Hyeon
Park, Jeongin
Lee, Soohyun
Kim, Dae Hyun
Shin, Sungbok
Seo, Jinwook
author_facet Jeon, Hyeon
Park, Jeongin
Lee, Soohyun
Kim, Dae Hyun
Shin, Sungbok
Seo, Jinwook
contents Selecting the appropriate dimensionality reduction (DR) technique and determining its optimal hyperparameter settings that maximize the accuracy of the output projections typically involves extensive trial and error, often resulting in unnecessary computational overhead. To address this challenge, we propose a dataset-adaptive approach to DR optimization guided by structural complexity metrics. These metrics quantify the intrinsic complexity of a dataset, predicting whether higher-dimensional spaces are necessary to represent it accurately. Since complex datasets are often inaccurately represented in two-dimensional projections, leveraging these metrics enables us to predict the maximum achievable accuracy of DR techniques for a given dataset, eliminating redundant trials in optimizing DR. We introduce the design and theoretical foundations of these structural complexity metrics. We quantitatively verify that our metrics effectively approximate the ground truth complexity of datasets and confirm their suitability for guiding dataset-adaptive DR workflow. Finally, we empirically show that our dataset-adaptive workflow significantly enhances the efficiency of DR optimization without compromising accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11984
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dataset-Adaptive Dimensionality Reduction
Jeon, Hyeon
Park, Jeongin
Lee, Soohyun
Kim, Dae Hyun
Shin, Sungbok
Seo, Jinwook
Human-Computer Interaction
Machine Learning
Selecting the appropriate dimensionality reduction (DR) technique and determining its optimal hyperparameter settings that maximize the accuracy of the output projections typically involves extensive trial and error, often resulting in unnecessary computational overhead. To address this challenge, we propose a dataset-adaptive approach to DR optimization guided by structural complexity metrics. These metrics quantify the intrinsic complexity of a dataset, predicting whether higher-dimensional spaces are necessary to represent it accurately. Since complex datasets are often inaccurately represented in two-dimensional projections, leveraging these metrics enables us to predict the maximum achievable accuracy of DR techniques for a given dataset, eliminating redundant trials in optimizing DR. We introduce the design and theoretical foundations of these structural complexity metrics. We quantitatively verify that our metrics effectively approximate the ground truth complexity of datasets and confirm their suitability for guiding dataset-adaptive DR workflow. Finally, we empirically show that our dataset-adaptive workflow significantly enhances the efficiency of DR optimization without compromising accuracy.
title Dataset-Adaptive Dimensionality Reduction
topic Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2507.11984