Salvato in:
Dettagli Bibliografici
Autore principale: Chang, Yuan-chin Ivan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2502.11036
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915156912504832
author Chang, Yuan-chin Ivan
author_facet Chang, Yuan-chin Ivan
contents Dimensionality reduction is a fundamental technique in machine learning and data analysis, enabling efficient representation and visualization of high-dimensional data. This paper explores five key methods: Principal Component Analysis (PCA), Kernel PCA (KPCA), Sparse Kernel PCA, t-Distributed Stochastic Neighbor Embedding (t-SNE), and Uniform Manifold Approximation and Projection (UMAP). PCA provides a linear approach to capturing variance, whereas KPCA and Sparse KPCA extend this concept to non-linear structures using kernel functions. Meanwhile, t-SNE and UMAP focus on preserving local relationships, making them effective for data visualization. Each method is examined in terms of its mathematical formulation, computational complexity, strengths, and limitations. The trade-offs between global structure preservation, computational efficiency, and interpretability are discussed to guide practitioners in selecting the appropriate technique based on their application needs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11036
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey: Potential Dimensionality Reduction Methods
Chang, Yuan-chin Ivan
Other Statistics
Dimensionality reduction is a fundamental technique in machine learning and data analysis, enabling efficient representation and visualization of high-dimensional data. This paper explores five key methods: Principal Component Analysis (PCA), Kernel PCA (KPCA), Sparse Kernel PCA, t-Distributed Stochastic Neighbor Embedding (t-SNE), and Uniform Manifold Approximation and Projection (UMAP). PCA provides a linear approach to capturing variance, whereas KPCA and Sparse KPCA extend this concept to non-linear structures using kernel functions. Meanwhile, t-SNE and UMAP focus on preserving local relationships, making them effective for data visualization. Each method is examined in terms of its mathematical formulation, computational complexity, strengths, and limitations. The trade-offs between global structure preservation, computational efficiency, and interpretability are discussed to guide practitioners in selecting the appropriate technique based on their application needs.
title A Survey: Potential Dimensionality Reduction Methods
topic Other Statistics
url https://arxiv.org/abs/2502.11036